LurkAPI
Docs · Scrape a page (Auto)

Scrape a page (Auto)

Any public page, escalating only as far as the site makes us. The price follows the strategy: 1–24 credits.

1 credit per call; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solvedMCP tool: web_scrape
GEThttps://api.lurkapi.com/v1/web/scrape

Example responseView as markdown

When to use this

Pass a url. Auto tries the strategies cheapest first, escalating only when a site walls it, and you pay for the one that got the page (the web scraping guide has details):

1. direct and static: a plain request, then a static ISP IP impersonating Chrome (TLS, HTTP/2 and headers). 1 credit. 2. browser: a real Chrome renders it, for JavaScript challenges (Cloudflare, AWS WAF, Akamai, PerimeterX) and pages that are only a JavaScript app. +3. 3. unblocker: a partner's browser on its own IPs, for sites that refuse ours (DataDome and the like). It returns the HTML the site sent, before JavaScript runs. +3. 4. residential: a residential IP, the last resort. +4. 5. browser_residential: our browser on a residential IP, the last resort for pages that need one. +13. 6. captcha: when a captcha still blocks our browser (Cloudflare Turnstile, DataDome's slider, reCAPTCHA), a solver clears it. +10, only when one is solved.

The price varies with the site, from 1 to 24 credits: billing in every answer says what this call cost and why. For a fixed price, call a strategy's own endpoint (/v1/web/fetch, /v1/web/render…); to cap it, chain the ones you accept with /v1/web/custom, or use render and proxy here: render=false&proxy=static never costs more than 1 credit; render=true always returns the page after JavaScript ran (never the unblocker).

via says which strategy delivered the page and attempts lists every one tried. A site that recently needed a stronger strategy starts there for 15 minutes, and once we get past a site's wall, later calls reuse that pass, often as a plain request for 1 credit. A call that every allowed strategy failed fails again at once, free, for 5 minutes; change render or proxy to try other strategies.

The site's own status comes back in status (a 404 page is charged like any page). When every allowed strategy is walled, or the site is down or answers 5xx, the call fails with 502 upstream_error and costs nothing; the message says why. Pages over 2 MB are cut (truncated). Use format=markdown for LLMs: links, headings, lists and tables survive, scripts and styles don't. Not cached: every call fetches the page. Calls take 1–3 s on most sites, 5–15 s when our browser renders, 15–35 s through the unblocker, and up to 85 s on the slowest (a captcha solved, a walled first visit retried).

Parameters

All parameters go in the query string.

ParameterDescription
url
stringrequired
The page to fetch: a full http(s) URL on a public host (ports 80 and 443 only).
Example
https://books.toscrape.com/
format
stringoptional
What content holds: the page's html, markdown (best for LLMs) or plain text. Non-HTML pages (JSON, XML, plain text) come back as they are.
One of
htmlmarkdowntext
Default
html
Example
markdown
render
stringoptional
auto: a browser (ours or the unblocker) only when the page needs one (+3 when it does). true: always our browser, and the page after JavaScript ran (+3). false: plain requests only.
One of
autotruefalse
Default
auto
proxy
stringoptional
auto: static ISP IPs and the unblocker; a residential IP only when they're refused (+4 when it is, +13 with our browser). static: never residential. residential: residential IPs only.
One of
autostaticresidential
Default
auto
country
stringoptional
Fetch from a residential IP in this country (ISO code like us, gb, de); implies proxy=residential (+4).
captcha
booleanoptional
true: solve a captcha that blocks the browser (+10, only when one is solved). false: never pay for a solve.
One of
truefalse
Default
true
wait_for
stringoptional
A CSS selector to wait for before returning, like .price or #reviews. Renders the page (+3).
wait
integeroptional
Milliseconds to let the page settle after it loads, up to 10000. Renders the page (+3).

Example request

Replace YOUR_API_KEY with your key, or set LURKAPI_KEY for the code. Get a free key.

Language

Example response

A real response from this endpoint, captured from the live API and trimmed to a couple of items. Strings over 96 characters (mostly signed media URLs) are cut short and end in ….

Show the example response (1 KB)
{
  "success": true,
  "credits_remaining": 9999,
  "credits_charged": 1,
  "status": 200,
  "url": "https://books.toscrape.com/",
  "content_type": "text/html",
  "content": "[Books to Scrape](https://books.toscrape.com/index.html) We love being s…",
  "truncated": false,
  "via": "direct",
  "captcha": null,
  "attempts": [
    {
      "via": "direct",
      "ms": 504,
      "status": 200,
      "wall": null,
      "error": null
    }
  ],
  "billing": {
    "mode": "auto",
    "credits": 1,
    "breakdown": {
      "base": 1
    },
    "note": "Auto: the price follows the strategy that got the page: 1 for a plain re…"
  }
}

Response fields

Every field of a successful response. [] marks a list: a[].b is the b of each item in a. nullable fields can be null; optional fields can be missing.

Show all 21 fields
FieldTypeDescription
successtrueAlways true here; errors have success: false.
credits_remainingnumberYour balance after this call. On anonymous playground calls: free tries left today.
credits_chargednumberCredits this call cost; 0 on free endpoints. On anonymous playground calls: tries used (1).
statusintegerThe page's HTTP status: 200, or the site's own 404, 410 and so on.
urlstringThe page's URL after redirects; the URL you asked for when via is unblocker, which doesn't say.
content_typestringnullableThe page's Content-Type, like text/html; charset=utf-8; null if the site sent none.
contentstringnullableThe page as format asks: HTML, markdown or plain text. Null for binary content (images, PDFs, archives).
truncatedbooleanThe page was over 2 MB and was cut there.
viastringThe strategy that got the page: direct or static (plain requests), browser (our browser), unblocker (a partner's browser and IPs), residential (a residential IP) or browser_residential (both).
captchastringnullableThe captcha solved to get the page (turnstile, datadome or recaptcha; charged); null if none was.
attemptsobject[]Every strategy tried, in order; the last one delivered the page.
attempts[].viastringThe strategy tried.
attempts[].msintegerHow long it took, in milliseconds.
attempts[].statusintegernullableThe HTTP status it got; null if it got no answer.
attempts[].wallstringnullableWhat walled it, like cloudflare challenge or datadome block; null if nothing did.
attempts[].errorstringnullableWhy it failed without an answer; null if it got one.
billingobjectWhat this call costs and why.
billing.modestringfixed: this endpoint's price. auto and custom: the price of the strategy that got the page.
billing.creditsintegerWhat this call costs: the sum of breakdown.
billing.breakdownobjectCredits by what they paid for: base, and on Auto or Custom render (a browser), premium (a residential IP) or premium_render (a browser on one); captcha when one was solved.
billing.notestringnullableWhy the price varies, on Auto and Custom; null on fixed-price endpoints.

Caching and freshness

Not cached: every call is answered fresh.

Errors

Errors return { success: false, error, code, docs } with the HTTP status below. Validation errors, 401/402/429 rejections and 5xx failures are free; not_found is charged (how charging works). All error codes.

StatusCodeMeaning
401missing_api_keyNo API key. Send it in the x-api-key header.
401invalid_api_keyThe key is unknown or was revoked.
402insufficient_creditsNot enough credits for this call. Buy a pack or wait for tomorrow's top-up.
403account_suspendedThis account is suspended. Contact support@lurkapi.com. Never charged.
405method_not_allowedEndpoints take GET with query params.
429rate_limitedToo many calls at once. Wait for the retry-after seconds, then retry.
500internal_errorSomething broke on our side. Retry; 5xx errors are free.
502upstream_errorThe platform didn't give a usable answer. Retry; 5xx errors are free.
503upstream_busyAll our connections to the platform are busy. Retry in a few seconds; 5xx errors are free.
504upstream_timeoutThe platform took over 30 seconds to answer (90 for web pages). Retry; timeouts are free.

Use it in Claude

Once LurkAPI is connected to Claude, this endpoint is the web_scrape tool. Ask in plain English, for example:

Ask Claude
Use LurkAPI's web_scrape with url "https://books.toscrape.com/", format "markdown" and summarize what you find.

Claude calls web_scrape with arguments like these, and each call costs 1 credit; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solved:

Tool arguments
{
  "url": "https://books.toscrape.com/",
  "format": "markdown"
}

Tool results skip nulls and empty lists to save tokens.