Find your way around the code
The whole package is about 3,500 lines of Python. It needs four libraries:
httpx, beautifulsoup4, playwright and Pillow.
Directorysrc/frankensurf/
- __init__.py makes
RuntimeandWebPolicyimportable - runtime.py the heart of it: settings, picking a tool, fetching, saving
- identity.py logins and saved secrets
- experimental.py the stealth browsers and per-site tool preferences
- provider_worker.py runs the stealth browsers in their own process
- search.py search engines
- cli.py the
frankensurfcommand - mcp_server.py the MCP server for AI apps
- commerce.py, classifieds.py, vehicles.py, ebay.py, gumtree.py, grays.py, allbids.py, depop.py site parsers
- __init__.py makes
Directoryscripts/ setup scripts and live tests
- …
Directorytests/ automated tests
- …
The main files
Section titled “The main files”| File | Lines | What’s in it |
|---|---|---|
runtime.py |
1,049 | WebPolicy (the settings), Runtime (read, extract, search, batch and friends), the tool order, fetching, the cache, saved copies and traces. |
identity.py |
565 | Registering logins, checking them, making sure only one process uses a login at a time, and secret storage. |
experimental.py |
193 | Checks whether the stealth browsers are installed, holds the per-site preferences, and runs a stealth fetch. |
provider_worker.py |
293 | A small separate program that runs one stealth fetch and prints the result as JSON. Also handles Carsales page clicks. |
search.py |
70 | Builds search requests and tidies the results for SearXNG, Bing and DuckDuckGo. |
cli.py |
153 | The command line. |
mcp_server.py |
89 | Seven MCP tools that call the same Runtime. |
The eight site parser files add up to about 1,070 lines. See Site parsers.
What happens when you call read
Section titled “What happens when you call read”Runtime.read(url, settings) 1. check the per-site preferred tool (unless you chose one) 2. check the URL is valid 3. if running as you: check the login and lock it 4. look for a saved copy (skipped when freshness is "now") 5. stop if this site already blocked us ← becomes a setting 6. make the list of tools to try 7. for each tool: fetch the page run the site parser, if any empty page that needs JavaScript? try the next tool save a copy, named by its hash download photos, if asked 8. return the result and receipt, save the trace