Skip to content

Python API

from frankensurf import Runtime, WebPolicy

Open it with async with. It looks after saved copies, the cache, logins and traces.

Runtime(
state_dir="state", # where saved copies, cache and traces go
steel_api_url=None, # your local Steel, or FRANKENSURF_STEEL_URL
local_cdp_url=None, # your Chrome; needs a registered login
concurrency=4, # how many pages at once in a batch
per_domain=1, # how many at once per site
domain_delay=0.25, # seconds between requests to one site
transport=None, # custom HTTP transport (for tests)
identity_registry=None, # or FRANKENSURF_IDENTITIES
)

For now Steel must run on your own machine.

Call Gives back Notes
await read(url, policy=None, provider=None, adapter=None) result Fetch a page.
await extract(url, adapter="html", policy=None, provider=None) result Fetch a page and run a site parser.
await search(query, source="searxng", limit=10, engine_config=None, policy=None) search results source is searxng, bing_rss or duckduckgo_html. Up to 100 results.
await batch(urls, policy=None, adapter=None) list of results Many pages at once, in the order given. Logged-in batches run one at a time.
await download_images(urls, policy=None) list of photos Downloads and checks each photo.
import_evidence(content, url, observed_at, …) result Save something you captured yourself, with the real time.
capabilities(domain=None) list Success and speed per site and tool, from past traces.
trace(trace_id) dict The full record of one request.
identity_status(identity_id=None) dict Whether your logins are healthy.

The settings for a request. Bad values are rejected straight away.

Setting Default What it does
freshness "now" now always fetches. hour, day and cached may reuse a saved copy.
provider None Force one tool: http, local, local_cdp, steel, camoufox, scrapling or scrapling_http.
render False Skip plain HTTP and use a browser.
include_images False Download the page’s photos.
max_images 12 Most photos to download (0 to 100).
timeout_seconds 25 How long to wait.
max_bytes 40 MB Biggest page to accept.
max_image_bytes 20 MB Biggest photo to accept.
allow_local_browser True Allow a browser on your machine.
allow_paid_fallbacks False Allow paid tools (for future paid plugins).
identity None Fetch as one of your logins.
wait_selector None Wait for this element before reading the page.
wait_state "attached" attached or visible.
settle_ms 400 Extra wait after the page is ready (up to 10 seconds).
navigation_page 1 Pages 1 to 3, for Carsales.

When something fails, you get a result with receipt["failure"] set to a code and a message, rather than an exception. See Error codes.

{
"url": "…",
"title": "…",
"text": "…", # readable text
"content": "…", # the raw page (left out by the CLI unless --raw)
"content_type": "text/html",
"headers": {…}, # never includes cookies
"structured": {…}, # what the site parser found
"image_urls": […],
"images": […], # downloaded photos, if asked
"field_status": {"availability": "unknown", "transaction_price": "unknown"},
"receipt": {…}, # see Receipts
}