Python API
from frankensurf import Runtime, WebPolicyRuntime
Section titled “Runtime”Open it with async with. It looks after saved copies, the cache, logins and
traces.
Runtime( state_dir="state", # where saved copies, cache and traces go steel_api_url=None, # your local Steel, or FRANKENSURF_STEEL_URL local_cdp_url=None, # your Chrome; needs a registered login concurrency=4, # how many pages at once in a batch per_domain=1, # how many at once per site domain_delay=0.25, # seconds between requests to one site transport=None, # custom HTTP transport (for tests) identity_registry=None, # or FRANKENSURF_IDENTITIES)For now Steel must run on your own machine.
What you can call
Section titled “What you can call”| Call | Gives back | Notes |
|---|---|---|
await read(url, policy=None, provider=None, adapter=None) |
result | Fetch a page. |
await extract(url, adapter="html", policy=None, provider=None) |
result | Fetch a page and run a site parser. |
await search(query, source="searxng", limit=10, engine_config=None, policy=None) |
search results | source is searxng, bing_rss or duckduckgo_html. Up to 100 results. |
await batch(urls, policy=None, adapter=None) |
list of results | Many pages at once, in the order given. Logged-in batches run one at a time. |
await download_images(urls, policy=None) |
list of photos | Downloads and checks each photo. |
import_evidence(content, url, observed_at, …) |
result | Save something you captured yourself, with the real time. |
capabilities(domain=None) |
list | Success and speed per site and tool, from past traces. |
trace(trace_id) |
dict | The full record of one request. |
identity_status(identity_id=None) |
dict | Whether your logins are healthy. |
WebPolicy
Section titled “WebPolicy”The settings for a request. Bad values are rejected straight away.
| Setting | Default | What it does |
|---|---|---|
freshness |
"now" |
now always fetches. hour, day and cached may reuse a saved copy. |
provider |
None |
Force one tool: http, local, local_cdp, steel, camoufox, scrapling or scrapling_http. |
render |
False |
Skip plain HTTP and use a browser. |
include_images |
False |
Download the page’s photos. |
max_images |
12 |
Most photos to download (0 to 100). |
timeout_seconds |
25 |
How long to wait. |
max_bytes |
40 MB | Biggest page to accept. |
max_image_bytes |
20 MB | Biggest photo to accept. |
allow_local_browser |
True |
Allow a browser on your machine. |
allow_paid_fallbacks |
False |
Allow paid tools (for future paid plugins). |
identity |
None |
Fetch as one of your logins. |
wait_selector |
None |
Wait for this element before reading the page. |
wait_state |
"attached" |
attached or visible. |
settle_ms |
400 |
Extra wait after the page is ready (up to 10 seconds). |
navigation_page |
1 |
Pages 1 to 3, for Carsales. |
Errors
Section titled “Errors”When something fails, you get a result with receipt["failure"] set to a code
and a message, rather than an exception. See Error codes.
What a result looks like
Section titled “What a result looks like”{ "url": "…", "title": "…", "text": "…", # readable text "content": "…", # the raw page (left out by the CLI unless --raw) "content_type": "text/html", "headers": {…}, # never includes cookies "structured": {…}, # what the site parser found "image_urls": […], "images": […], # downloaded photos, if asked "field_status": {"availability": "unknown", "transaction_price": "unknown"}, "receipt": {…}, # see Receipts}