LurkAPI
Docs · Scraping any web page

Scraping any web page

Any public URL as HTML, markdown or text, past bot walls and captchas. Let Auto find the cheapest way through, from 1 credit, or pick the way and its price yourself.

Try it liveView as markdown

What it does

GET /v1/web/scrape fetches the page at url and returns it in content, in the format you ask for: html (the default), markdown or text. Use markdown for LLMs: headings, links, lists and tables survive; scripts and styles don't. JSON, XML and plain-text pages come back as they are; binary files (images, PDFs) come back with content: null.

When a site walls the request (a JavaScript challenge, an IP block, a captcha), the call escalates, cheapest way first, until one gets through. You pay only for the one that delivered the page. Most sites need none of it: 1 credit, 1–3 seconds.

status is the site's own HTTP status and url the page's URL after redirects. In Claude and other MCP clients this is the web_scrape tool. Every parameter and field.

Auto, a fixed price, or your own steps

Seven endpoints share one engine and one response. Auto picks the way for you; the fixed-price ones each run one way; Custom runs the ways you list.

EndpointRunsCredits
Auto /v1/web/scrapethe whole ladder below, cheapest first1–24, set by the step that got the page
Fetch a page (plain request) /v1/web/fetchdirect then static1 credit
Render a page (browser) /v1/web/renderbrowser4 credits; 14 if a captcha is solved
Unblock a page (partner browser) /v1/web/unblockunblocker4 credits
Fetch a page from a residential IP /v1/web/residentialresidential5 credits
Render a page from a residential IP /v1/web/render-residentialbrowser_residential14 credits; 24 if a captcha is solved
Custom /v1/web/customthe steps you list in steps, in your orderset by the step that got the page, at its fixed price

Every answer says what it cost in billing: mode (auto, fixed or custom), credits (the same as credits_charged), a breakdown (base, render, premium, premium_render, captcha) and, on Auto and Custom, a note on why the price varies. The fixed-price endpoints and Custom solve a captcha (+10) only with captcha=true, so their price stays what you chose unless you allow one.

Custom is for when you know a site: steps=fetch,unblock never costs more than 4 credits and never runs JavaScript; steps=render,render-residential always returns the page as rendered. It skips a step a wall has already ruled out, the way Auto does.

Try it live

Pick an endpoint and run it: real data, no signup, 5 free requests a day.

Endpoint

GET/v1/web/scrape

1 credit per request; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solved
Examples
url

The page to fetch: a full http(s) URL on a public host (ports 80 and 443 only).

format

What content holds: the page's html, markdown (best for LLMs) or plain text. Non-HTML pages (JSON, XML, plain text) come back as they are.

More options (6)
render

auto: a browser (ours or the unblocker) only when the page needs one (+3 when it does). true: always our browser, and the page after JavaScript ran (+3). false: plain requests only.

proxy

auto: static ISP IPs and the unblocker; a residential IP only when they're refused (+4 when it is, +13 with our browser). static: never residential. residential: residential IPs only.

country

Fetch from a residential IP in this country (ISO code like us, gb, de); implies proxy=residential (+4).

captcha

true: solve a captcha that blocks the browser (+10, only when one is solved). false: never pay for a solve.

wait_for

A CSS selector to wait for before returning, like .price or #reviews. Renders the page (+3).

wait

Milliseconds to let the page settle after it loads, up to 10000. Renders the page (+3).

5 free tries a day, no signup

Press Run to fetch live data

Results come back as a table you can download as CSV, or as the raw JSON.

How it gets the page

An Auto call climbs this ladder and stops at the first step that gets the page:

StepHowForExtra credits
directA plain requestMost sitesnone
staticA static ISP IP with a real Chrome fingerprint (TLS, HTTP/2 and headers)Sites that refuse a datacenter IP or a client that isn't a browsernone
browserA real Chrome on a static ISP IPJavaScript challenges (Cloudflare, AWS WAF, Akamai, PerimeterX) and pages that are only a JavaScript app+3
unblockerA partner's browser on its own IPsSites that refuse our IPs (DataDome and the like). It returns the HTML the site sent, before JavaScript runs, and the URL you asked for, not where redirects ended.+3
residentialA residential IPThe last resort, when the steps above are refused+4
browser_residentialOur browser on a residential IPThe last resort for pages that need a browser+13
captchaA captcha solver, inside our browserA Cloudflare Turnstile, DataDome slider or reCAPTCHA our browser can't click past+10, only when one is solved

What stopped the last step decides the next: a JavaScript challenge skips the remaining plain requests for a browser, an IP block (DataDome, a bare 403 or 429) skips our IPs for the partner's browser, and an empty JavaScript app goes to our browser. So a Cloudflare-challenged site typically goes direct → static → browser, and a DataDome site direct → static → unblocker.

via names the step that delivered the page, and captcha the captcha solved for it (turnstile, datadome or recaptcha), or null. attempts lists every step tried, in order: how long it took (ms), the HTTP status it got, and what walled it (wall, like cloudflare challenge or datadome block) or why it failed (error).

A DataDome site's first call: Yelp, in production (trimmed)
{
  "credits_charged": 4,
  "billing": { "mode": "auto", "credits": 4, "breakdown": { "base": 1, "render": 3 }, "note": "Auto: the price follows the strategy that got the page…" },
  "via": "unblocker",
  "captcha": null,
  "attempts": [
    { "via": "direct", "ms": 43, "status": 403, "wall": "datadome block", "error": null },
    { "via": "static", "ms": 802, "status": 403, "wall": "datadome block", "error": null },
    { "via": "unblocker", "ms": 17651, "status": 200, "wall": null, "error": null }
  ]
}

Credits

  • 1 credit for every page delivered. Paid steps add to it: +3 when a browser renders it (ours or the partner's), +4 for a residential IP, +13 for our browser on one, +10 when a captcha is solved. A call costs 1–24 credits.
  • You pay for the step that got the page. A paid step is held from your balance before it runs and handed back if it doesn't deliver, so a render that hit a wall followed by a residential IP that got the page costs 5, not 8.
  • Failures are free: walled at every step, the site down, the site's own 5xx, or no answer in time.
  • A page is a page: the site's own 404 or 410 comes back with that status and is charged.
  • Paid steps need the credits when they start. If your balance can't cover a step the site needs, the call stops with 402 insufficient_credits, free. A captcha solve is the exception: our browser still renders, without the solver.
The siteUsually delivered byCredits
Most sitesdirect or static1
A JavaScript challenge, or a page that's only a JavaScript appbrowser4
Refuses our IPs (DataDome and the like)unblocker4
A captcha our browser can't click pastbrowser and a solve14
Only a residential IP gets throughresidential or browser_residential5 or 14
Only a residential IP gets through, behind a captchabrowser_residential and a solve24

Controlling cost and behaviour

On Auto, the defaults escalate as far as the site makes them. These parameters cap or force the ladder (or use a fixed-price endpoint, or Custom):

ParameterValuesWhat it does
renderauto (default), true, falseauto: a browser only when the page needs one. true: always our browser, so content is the page after JavaScript ran (never the partner's browser, which returns the HTML as sent). false: plain requests only.
proxyauto (default), static, residentialauto: static ISP IPs and the partner's browser; a residential IP only when they're refused. static: never residential. residential: residential IPs only.
countryA two-letter code: us, gb, de…Fetch from a residential IP in that country. Implies proxy=residential.
captchatrue (default), falsefalse: never pay for a solve.
wait_forA CSS selector, like .priceWait until it's on the page before returning. Implies render=true.
waitMilliseconds, up to 10000Let the page settle this long after it loads. Implies render=true.

What each choice can cost, in credits:

ParametersSteps allowedLeastMost
The defaultsall124
captcha=falseall, without solves114
proxy=staticdirect, static, browser, unblocker114
proxy=static&captcha=falsethe same, without solves14
render=falsedirect, static, residential15
render=false&proxy=staticdirect, static11
render=true, wait_for or waitbrowser, browser_residential424
proxy=residential or countryresidential, browser_residential524

So render=false&proxy=static never costs more than 1 credit, and render=true always returns the page after JavaScript ran. A narrower ladder can fail where the defaults would get through; after render=false, proxy=static or captcha=false, the error says which to relax.

Speed

  • Most pages: 1–3 s.
  • Rendered by our browser: 5–15 s.
  • Through the partner's browser: 15–35 s.
  • Up to 85 s for a solved captcha or the slowest sites. Give your HTTP client a timeout of at least 90 s; the API answers 504 upstream_timeout by then.

Repeat calls get faster:

  • A site that needed a stronger step starts there for the next 15 minutes (for calls allowed the same steps), instead of hitting the same walls again.
  • Once our browser gets past a site's wall, later calls to that site reuse the pass for up to 30 minutes: often as a plain request (1 credit, a few seconds), else as a render that skips the wall, so a captcha isn't solved, or paid for, twice.

Failures

A call that gets no page fails with 502 upstream_error and costs nothing. The error says why:

  • Walled at every step it was allowed. The message names each step and its wall, and the options that might get through (Allowing render=auto and proxy=auto may get through.). The same call (same url and options) then fails at once for 5 minutes; change render or proxy to try other steps.
  • The site is down: it doesn't resolve, its server is down behind Cloudflare (521–530), it refuses connections, its certificate is invalid, or no network got an answer. Every call to that host then fails at once for 3 minutes.
  • The site answered 5xx, or didn't answer in time: retry.
StatusCodeWhen
400invalid_paramsA missing or invalid parameter, or a url we won't fetch: not http(s), a port other than 80 or 443, credentials in it, or a host that isn't public (localhost, a private IP, or a name that resolves to one). issues names it.
502upstream_errorNo page: walled at every step allowed, the site down, its 5xx, or no answer. Free.
504upstream_timeoutThe call ran past 90 s. Free; rare, since the ladder stops at 85 s.

Limits and good use

  • Public pages only. The endpoint takes no cookies or headers from you, so nothing behind a login.
  • Public hosts on ports 80 and 443, in URLs up to 2,048 characters.
  • Encode url like any query value: curl -G --data-urlencode "url=…" does it for you.
  • Pages over 2 MB are cut there, and truncated is true.
  • Not cached. Every call fetches the page and costs credits; keep your own copy if you need a page twice.
  • Use format=markdown for LLMs. It's much smaller than the HTML and keeps the structure a model needs.

Examples

A page as markdown

Language

A walled site

Nothing to set: the defaults escalate until the page comes through. A URL with its own query string has to be encoded:

curl
curl -G "https://api.lurkapi.com/v1/web/scrape" \
  --data-urlencode "url=https://www.tripadvisor.com/Search?q=lisbon" \
  -d format=markdown \
  -H "x-api-key: YOUR_API_KEY"

Always 1 credit

Plain requests only: a walled site fails, free, instead of escalating.

curl
curl -G "https://api.lurkapi.com/v1/web/fetch" \
  --data-urlencode "url=https://news.ycombinator.com/" \
  -d format=text \
  -H "x-api-key: YOUR_API_KEY"

Your own steps, at most 4 credits

A plain request first, then the partner's browser if the site walls it:

curl
curl -G "https://api.lurkapi.com/v1/web/custom" \
  --data-urlencode "url=https://www.yelp.com/biz/the-house-san-francisco" \
  -d steps=fetch,unblock \
  -d format=markdown \
  -H "x-api-key: YOUR_API_KEY"

FAQ

Why did I pay 14 credits?

Our browser rendered the page (+3) and a captcha solver cleared a captcha it couldn't click past (+10); captcha in the response names it. Pass captcha=false to never pay for a solve: the call then tries the other steps and fails, free, if none gets through.

Why is a site “down”?

The site itself isn't answering: it doesn't resolve, its server is down behind Cloudflare, it refuses connections, its certificate is invalid, or no network got an answer. No other IP or browser changes that, so calls to it fail at once, free, for 3 minutes; then we try it again.

Can I scrape logged-in pages?

No. It fetches public pages only and takes no cookies or headers from you. A page behind a login comes back as whatever the site shows a visitor who isn't signed in, usually its login page.

How is this different from the platform endpoints?

The platform endpoints return structured JSON fields at a listed price per call, and are cached. Web scrape returns any page as HTML, markdown or text, fresh every time, and costs what the site made it take (1–24 credits). If a platform endpoint covers what you need, use it.