Docs · Scrape a page (Auto)
Scrape a page (Auto)
Any public page, escalating only as far as the site makes us. The price follows the strategy: 1–24 credits.
https://api.lurkapi.com/v1/web/scrapeWhen to use this
Pass a url. Auto tries the strategies cheapest first, escalating only when a site walls it, and you pay for the one that got the page (the web scraping guide has details):
1. direct and static: a plain request, then a static ISP IP impersonating Chrome (TLS, HTTP/2 and headers). 1 credit. 2. browser: a real Chrome renders it, for JavaScript challenges (Cloudflare, AWS WAF, Akamai, PerimeterX) and pages that are only a JavaScript app. +3. 3. unblocker: a partner's browser on its own IPs, for sites that refuse ours (DataDome and the like). It returns the HTML the site sent, before JavaScript runs. +3. 4. residential: a residential IP, the last resort. +4. 5. browser_residential: our browser on a residential IP, the last resort for pages that need one. +13. 6. captcha: when a captcha still blocks our browser (Cloudflare Turnstile, DataDome's slider, reCAPTCHA), a solver clears it. +10, only when one is solved.
The price varies with the site, from 1 to 24 credits: billing in every answer says what this call cost and why. For a fixed price, call a strategy's own endpoint (/v1/web/fetch, /v1/web/render…); to cap it, chain the ones you accept with /v1/web/custom, or use render and proxy here: render=false&proxy=static never costs more than 1 credit; render=true always returns the page after JavaScript ran (never the unblocker).
via says which strategy delivered the page and attempts lists every one tried. A site that recently needed a stronger strategy starts there for 15 minutes, and once we get past a site's wall, later calls reuse that pass, often as a plain request for 1 credit. A call that every allowed strategy failed fails again at once, free, for 5 minutes; change render or proxy to try other strategies.
The site's own status comes back in status (a 404 page is charged like any page). When every allowed strategy is walled, or the site is down or answers 5xx, the call fails with 502 upstream_error and costs nothing; the message says why. Pages over 2 MB are cut (truncated). Use format=markdown for LLMs: links, headings, lists and tables survive, scripts and styles don't. Not cached: every call fetches the page. Calls take 1–3 s on most sites, 5–15 s when our browser renders, 15–35 s through the unblocker, and up to 85 s on the slowest (a captcha solved, a walled first visit retried).
Parameters
All parameters go in the query string.
| Parameter | Description |
|---|---|
urlstringrequired | The page to fetch: a full http(s) URL on a public host (ports 80 and 443 only).
|
formatstringoptional | What content holds: the page's html, markdown (best for LLMs) or plain text. Non-HTML pages (JSON, XML, plain text) come back as they are.
|
renderstringoptional | auto: a browser (ours or the unblocker) only when the page needs one (+3 when it does). true: always our browser, and the page after JavaScript ran (+3). false: plain requests only.
|
proxystringoptional | auto: static ISP IPs and the unblocker; a residential IP only when they're refused (+4 when it is, +13 with our browser). static: never residential. residential: residential IPs only.
|
countrystringoptional | Fetch from a residential IP in this country (ISO code like us, gb, de); implies proxy=residential (+4). |
captchabooleanoptional | true: solve a captcha that blocks the browser (+10, only when one is solved). false: never pay for a solve.
|
wait_forstringoptional | A CSS selector to wait for before returning, like .price or #reviews. Renders the page (+3). |
waitintegeroptional | Milliseconds to let the page settle after it loads, up to 10000. Renders the page (+3). |
Example request
Replace YOUR_API_KEY with your key, or set LURKAPI_KEY for the code. Get a free key.
curl "https://api.lurkapi.com/v1/web/scrape?url=https%3A%2F%2Fbooks.toscrape.com%2F&format=markdown" \
-H "x-api-key: YOUR_API_KEY"const params = new URLSearchParams({
url: "https://books.toscrape.com/",
format: "markdown",
});
const res = await fetch(`https://api.lurkapi.com/v1/web/scrape?${params}`, {
headers: { "x-api-key": process.env.LURKAPI_KEY },
});
const data = await res.json();
if (!data.success) throw new Error(`${data.code}: ${data.error}`);
console.log(data.content);import os
import requests
res = requests.get(
"https://api.lurkapi.com/v1/web/scrape",
params={
"url": "https://books.toscrape.com/",
"format": "markdown",
},
headers={"x-api-key": os.environ["LURKAPI_KEY"]},
timeout=95,
)
data = res.json()
if not data["success"]:
raise RuntimeError(f"{data['code']}: {data['error']}")
print(data["content"])Example response
A real response from this endpoint, captured from the live API and trimmed to a couple of items. Strings over 96 characters (mostly signed media URLs) are cut short and end in ….
Show the example response (1 KB)Hide
{
"success": true,
"credits_remaining": 9999,
"credits_charged": 1,
"status": 200,
"url": "https://books.toscrape.com/",
"content_type": "text/html",
"content": "[Books to Scrape](https://books.toscrape.com/index.html) We love being s…",
"truncated": false,
"via": "direct",
"captcha": null,
"attempts": [
{
"via": "direct",
"ms": 504,
"status": 200,
"wall": null,
"error": null
}
],
"billing": {
"mode": "auto",
"credits": 1,
"breakdown": {
"base": 1
},
"note": "Auto: the price follows the strategy that got the page: 1 for a plain re…"
}
}Response fields
Every field of a successful response. [] marks a list: a[].b is the b of each item in a. nullable fields can be null; optional fields can be missing.
Show all 21 fieldsHide
| Field | Type | Description |
|---|---|---|
| success | true | Always true here; errors have success: false. |
| credits_remaining | number | Your balance after this call. On anonymous playground calls: free tries left today. |
| credits_charged | number | Credits this call cost; 0 on free endpoints. On anonymous playground calls: tries used (1). |
| status | integer | The page's HTTP status: 200, or the site's own 404, 410 and so on. |
| url | string | The page's URL after redirects; the URL you asked for when via is unblocker, which doesn't say. |
| content_type | stringnullable | The page's Content-Type, like text/html; charset=utf-8; null if the site sent none. |
| content | stringnullable | The page as format asks: HTML, markdown or plain text. Null for binary content (images, PDFs, archives). |
| truncated | boolean | The page was over 2 MB and was cut there. |
| via | string | The strategy that got the page: direct or static (plain requests), browser (our browser), unblocker (a partner's browser and IPs), residential (a residential IP) or browser_residential (both). |
| captcha | stringnullable | The captcha solved to get the page (turnstile, datadome or recaptcha; charged); null if none was. |
| attempts | object[] | Every strategy tried, in order; the last one delivered the page. |
| attempts[].via | string | The strategy tried. |
| attempts[].ms | integer | How long it took, in milliseconds. |
| attempts[].status | integernullable | The HTTP status it got; null if it got no answer. |
| attempts[].wall | stringnullable | What walled it, like cloudflare challenge or datadome block; null if nothing did. |
| attempts[].error | stringnullable | Why it failed without an answer; null if it got one. |
| billing | object | What this call costs and why. |
| billing.mode | string | fixed: this endpoint's price. auto and custom: the price of the strategy that got the page. |
| billing.credits | integer | What this call costs: the sum of breakdown. |
| billing.breakdown | object | Credits by what they paid for: base, and on Auto or Custom render (a browser), premium (a residential IP) or premium_render (a browser on one); captcha when one was solved. |
| billing.note | stringnullable | Why the price varies, on Auto and Custom; null on fixed-price endpoints. |
Caching and freshness
Not cached: every call is answered fresh.
Errors
Errors return { success: false, error, code, docs } with the HTTP status below. Validation errors, 401/402/429 rejections and 5xx failures are free; not_found is charged (how charging works). All error codes.
| Status | Code | Meaning |
|---|---|---|
| 401 | missing_api_key | No API key. Send it in the x-api-key header. |
| 401 | invalid_api_key | The key is unknown or was revoked. |
| 402 | insufficient_credits | Not enough credits for this call. Buy a pack or wait for tomorrow's top-up. |
| 403 | account_suspended | This account is suspended. Contact support@lurkapi.com. Never charged. |
| 405 | method_not_allowed | Endpoints take GET with query params. |
| 429 | rate_limited | Too many calls at once. Wait for the retry-after seconds, then retry. |
| 500 | internal_error | Something broke on our side. Retry; 5xx errors are free. |
| 502 | upstream_error | The platform didn't give a usable answer. Retry; 5xx errors are free. |
| 503 | upstream_busy | All our connections to the platform are busy. Retry in a few seconds; 5xx errors are free. |
| 504 | upstream_timeout | The platform took over 30 seconds to answer (90 for web pages). Retry; timeouts are free. |
Use it in Claude
Once LurkAPI is connected to Claude, this endpoint is the web_scrape tool. Ask in plain English, for example:
Use LurkAPI's web_scrape with url "https://books.toscrape.com/", format "markdown" and summarize what you find.Claude calls web_scrape with arguments like these, and each call costs 1 credit; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solved:
{
"url": "https://books.toscrape.com/",
"format": "markdown"
}Tool results skip nulls and empty lists to save tokens.