Site parsers
A site parser turns a fetched page into clean data: title, price, photos, next
page and so on. In the code they’re called adapters. Use one with
extract(url, adapter), --adapter on the command line, or adapter in MCP.
Rules every parser follows
Section titled “Rules every parser follows”- Right item. The listing ID on the page must match the URL you asked for. If not, it fails instead of returning the wrong thing.
- Only the item’s photos. No logos, ads, recommendations or duplicates. If the
photo count doesn’t add up, it fails with
SCHEMA_CHANGED. - Claims stay claims. “Buy now” buttons and status labels are recorded as what the page says, not as proof the item is available.
- Prices stay separate. Asking price, current bid and sale price are never mixed up.
- Careful paging. A “next page” link is only given when it keeps your search and filters. Reaching the last page doesn’t claim you’ve seen everything.
General parsers
Section titled “General parsers”| Parser | For |
|---|---|
html |
Any page: title, text, structured data and image links (the default) |
json |
JSON responses |
shopify_product, shopify_catalogue |
Shopify stores |
Australian marketplaces
Section titled “Australian marketplaces”| Site | Search results | Single listing | Photo gallery |
|---|---|---|---|
| Gumtree | gumtree_search |
gumtree_listing |
in listing |
| eBay AU | ebay_search |
ebay_listing |
in listing |
| Carsales | carsales_search |
carsales_listing |
carsales_gallery |
| Bikesales | bikesales_search¹ |
bikesales_listing |
bikesales_gallery |
| Depop | depop_search, depop_catalogue |
depop_listing |
in listing |
| Cash Converters | cashconverters_catalogue |
cashconverters_listing |
in listing |
| Trading Post | tradingpost_search |
tradingpost_listing |
in listing |
| Grays | grays_search |
grays_listing |
in listing |
| Lloyds | lloyds_search |
lloyds_listing |
in listing |
| ALLBIDS | allbids_catalogue |
allbids_listing |
in listing |
¹ Works from Python, but isn’t a command-line option yet.
Adding a parser
Section titled “Adding a parser”Today a parser is a function in one of the site files, and parse_content in
runtime.py calls it by name. With plugins, a parser becomes a self-contained
package: a name, the URLs it handles, its checks, and sample pages to test
against. Later, Frankensurf will be able to write new parsers itself from a
sample page, and only use them once they pass their tests.