# Scrape a page (Auto)

> Any public page, escalating only as far as the site makes us. The price follows the strategy: 1–24 credits.

- **Request:** `GET https://api.lurkapi.com/v1/web/scrape`
- **Auth:** `x-api-key` header ([get a free key](https://lurkapi.com/login?next=/dashboard))
- **Cost:** 1 credit per call; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solved. Validation errors, `401`/`402`/`429` rejections and `5xx` failures are free; `not_found` is charged ([how charging works](https://lurkapi.com/docs.md#credits)).
- **MCP tool:** `web_scrape` on `https://api.lurkapi.com/mcp`
- **Try it live:** https://lurkapi.com/docs/web/scrape#try (no signup)
- **Platform:** [Web](https://lurkapi.com/docs/web.md) (any public URL)
- **Web page:** https://lurkapi.com/docs/web/scrape

## When to use this

Pass a `url`. Auto tries the strategies cheapest first, escalating only when a site walls it, and you pay for the one that got the page (the [web scraping guide](https://lurkapi.com/docs/web-scraping.md) has details):

1. **direct** and **static**: a plain request, then a static ISP IP impersonating Chrome (TLS, HTTP/2 and headers). 1 credit.
2. **browser**: a real Chrome renders it, for JavaScript challenges (Cloudflare, AWS WAF, Akamai, PerimeterX) and pages that are only a JavaScript app. +3.
3. **unblocker**: a partner's browser on its own IPs, for sites that refuse ours (DataDome and the like). It returns the HTML the site sent, before JavaScript runs. +3.
4. **residential**: a residential IP, the last resort. +4.
5. **browser_residential**: our browser on a residential IP, the last resort for pages that need one. +13.
6. **captcha**: when a captcha still blocks our browser (Cloudflare Turnstile, DataDome's slider, reCAPTCHA), a solver clears it. +10, only when one is solved.

**The price varies with the site**, from 1 to 24 credits: `billing` in every answer says what this call cost and why. For a fixed price, call a strategy's own endpoint (`/v1/web/fetch`, `/v1/web/render`…); to cap it, chain the ones you accept with `/v1/web/custom`, or use `render` and `proxy` here: `render=false&proxy=static` never costs more than 1 credit; `render=true` always returns the page after JavaScript ran (never the unblocker).

`via` says which strategy delivered the page and `attempts` lists every one tried. A site that recently needed a stronger strategy starts there for 15 minutes, and once we get past a site's wall, later calls reuse that pass, often as a plain request for 1 credit. A call that every allowed strategy failed fails again at once, free, for 5 minutes; change `render` or `proxy` to try other strategies.

The site's own status comes back in `status` (a 404 page is charged like any page). When every allowed strategy is walled, or the site is down or answers 5xx, the call fails with 502 `upstream_error` and costs nothing; the message says why. Pages over 2 MB are cut (`truncated`). Use `format=markdown` for LLMs: links, headings, lists and tables survive, scripts and styles don't. Not cached: every call fetches the page. Calls take 1–3 s on most sites, 5–15 s when our browser renders, 15–35 s through the unblocker, and up to 85 s on the slowest (a captcha solved, a walled first visit retried).

## Parameters

All parameters go in the query string.

| Name | Type | Required | Default | Allowed values | Description | Example |
| --- | --- | --- | --- | --- | --- | --- |
| `url` | string | yes |  |  | The page to fetch: a full http(s) URL on a public host (ports 80 and 443 only). | `https://books.toscrape.com/` |
| `format` | string | no | `html` | `html`, `markdown`, `text` | What `content` holds: the page's `html`, `markdown` (best for LLMs) or plain `text`. Non-HTML pages (JSON, XML, plain text) come back as they are. | `markdown` |
| `render` | string | no | `auto` | `auto`, `true`, `false` | `auto`: a browser (ours or the unblocker) only when the page needs one (+3 when it does). `true`: always our browser, and the page after JavaScript ran (+3). `false`: plain requests only. |  |
| `proxy` | string | no | `auto` | `auto`, `static`, `residential` | `auto`: static ISP IPs and the unblocker; a residential IP only when they're refused (+4 when it is, +13 with our browser). `static`: never residential. `residential`: residential IPs only. |  |
| `country` | string | no |  |  | Fetch from a residential IP in this country (ISO code like `us`, `gb`, `de`); implies `proxy=residential` (+4). |  |
| `captcha` | boolean | no | `true` | `true`, `false` | `true`: solve a captcha that blocks the browser (+10, only when one is solved). `false`: never pay for a solve. |  |
| `wait_for` | string | no |  |  | A CSS selector to wait for before returning, like `.price` or `#reviews`. Renders the page (+3). |  |
| `wait` | integer | no |  |  | Milliseconds to let the page settle after it loads, up to 10000. Renders the page (+3). |  |

## Example request

Replace `YOUR_API_KEY` with your key, or set `LURKAPI_KEY` for the code. [Get a free key](https://lurkapi.com/login?next=/dashboard).

curl:

```bash
curl "https://api.lurkapi.com/v1/web/scrape?url=https%3A%2F%2Fbooks.toscrape.com%2F&format=markdown" \
  -H "x-api-key: YOUR_API_KEY"
```

JavaScript:

```js
const params = new URLSearchParams({
  url: "https://books.toscrape.com/",
  format: "markdown",
});
const res = await fetch(`https://api.lurkapi.com/v1/web/scrape?${params}`, {
  headers: { "x-api-key": process.env.LURKAPI_KEY },
});
const data = await res.json();
if (!data.success) throw new Error(`${data.code}: ${data.error}`);
console.log(data.content);
```

Python:

```python
import os
import requests

res = requests.get(
    "https://api.lurkapi.com/v1/web/scrape",
    params={
        "url": "https://books.toscrape.com/",
        "format": "markdown",
    },
    headers={"x-api-key": os.environ["LURKAPI_KEY"]},
    timeout=95,
)
data = res.json()
if not data["success"]:
    raise RuntimeError(f"{data['code']}: {data['error']}")
print(data["content"])
```

## Example response

A real response from this endpoint, captured from the live API and trimmed to a couple of items. Strings over 96 characters (mostly signed media URLs) are cut short and end in `…`.

200 OK (application/json):

```json
{
  "success": true,
  "credits_remaining": 9999,
  "credits_charged": 1,
  "status": 200,
  "url": "https://books.toscrape.com/",
  "content_type": "text/html",
  "content": "[Books to Scrape](https://books.toscrape.com/index.html) We love being s…",
  "truncated": false,
  "via": "direct",
  "captcha": null,
  "attempts": [
    {
      "via": "direct",
      "ms": 504,
      "status": 200,
      "wall": null,
      "error": null
    }
  ],
  "billing": {
    "mode": "auto",
    "credits": 1,
    "breakdown": {
      "base": 1
    },
    "note": "Auto: the price follows the strategy that got the page: 1 for a plain re…"
  }
}
```

## Response fields

Every field of a successful response. `[]` marks a list: `a[].b` is the `b` of each item in `a`. **nullable** fields can be `null`; **optional** fields can be missing.

| Field | Type | Description |
| --- | --- | --- |
| `success` | `true` | Always true here; errors have `success: false`. |
| `credits_remaining` | `number` | Your balance after this call. On anonymous playground calls: free tries left today. |
| `credits_charged` | `number` | Credits this call cost; 0 on free endpoints. On anonymous playground calls: tries used (1). |
| `status` | `integer` | The page's HTTP status: 200, or the site's own 404, 410 and so on. |
| `url` | `string` | The page's URL after redirects; the URL you asked for when `via` is `unblocker`, which doesn't say. |
| `content_type` | `string`, nullable | The page's Content-Type, like `text/html; charset=utf-8`; null if the site sent none. |
| `content` | `string`, nullable | The page as `format` asks: HTML, markdown or plain text. Null for binary content (images, PDFs, archives). |
| `truncated` | `boolean` | The page was over 2 MB and was cut there. |
| `via` | `string` | The strategy that got the page: `direct` or `static` (plain requests), `browser` (our browser), `unblocker` (a partner's browser and IPs), `residential` (a residential IP) or `browser_residential` (both). |
| `captcha` | `string`, nullable | The captcha solved to get the page (`turnstile`, `datadome` or `recaptcha`; charged); null if none was. |
| `attempts` | `object[]` | Every strategy tried, in order; the last one delivered the page. |
| `attempts[].via` | `string` | The strategy tried. |
| `attempts[].ms` | `integer` | How long it took, in milliseconds. |
| `attempts[].status` | `integer`, nullable | The HTTP status it got; null if it got no answer. |
| `attempts[].wall` | `string`, nullable | What walled it, like `cloudflare challenge` or `datadome block`; null if nothing did. |
| `attempts[].error` | `string`, nullable | Why it failed without an answer; null if it got one. |
| `billing` | `object` | What this call costs and why. |
| `billing.mode` | `string` | `fixed`: this endpoint's price. `auto` and `custom`: the price of the strategy that got the page. |
| `billing.credits` | `integer` | What this call costs: the sum of `breakdown`. |
| `billing.breakdown` | `object` | Credits by what they paid for: `base`, and on Auto or Custom `render` (a browser), `premium` (a residential IP) or `premium_render` (a browser on one); `captcha` when one was solved. |
| `billing.note` | `string`, nullable | Why the price varies, on Auto and Custom; null on fixed-price endpoints. |

## Caching and freshness

Not cached: every call is answered fresh.

## Errors

Errors return `{ success: false, error, code, docs }` with the HTTP status below. Validation errors, `401`/`402`/`429` rejections and `5xx` failures are free; `not_found` is charged ([how charging works](https://lurkapi.com/docs.md#credits)). [All error codes](https://lurkapi.com/docs.md#errors).

| Status | Code | Meaning |
| --- | --- | --- |
| 401 | `missing_api_key` | No API key. Send it in the `x-api-key` header. |
| 401 | `invalid_api_key` | The key is unknown or was revoked. |
| 402 | `insufficient_credits` | Not enough credits for this call. Buy a pack or wait for tomorrow's top-up. |
| 403 | `account_suspended` | This account is suspended. Contact support@lurkapi.com. Never charged. |
| 405 | `method_not_allowed` | Endpoints take GET with query params. |
| 429 | `rate_limited` | Too many calls at once. Wait for the `retry-after` seconds, then retry. |
| 500 | `internal_error` | Something broke on our side. Retry; 5xx errors are free. |
| 502 | `upstream_error` | The platform didn't give a usable answer. Retry; 5xx errors are free. |
| 503 | `upstream_busy` | All our connections to the platform are busy. Retry in a few seconds; 5xx errors are free. |
| 504 | `upstream_timeout` | The platform took over 30 seconds to answer (90 for web pages). Retry; timeouts are free. |

## Use it in Claude

Once LurkAPI is [connected to Claude](https://lurkapi.com/docs.md#claude), this endpoint is the `web_scrape` tool. Ask in plain English, for example:

Ask Claude:

```text
Use LurkAPI's web_scrape with url "https://books.toscrape.com/", format "markdown" and summarize what you find.
```

Claude calls `web_scrape` with arguments like these, and each call costs 1 credit; +3 if a browser renders it, +4 if it needs a residential IP, +13 if a browser renders it on a residential IP, +10 if a captcha is solved:

Tool arguments:

```json
{
  "url": "https://books.toscrape.com/",
  "format": "markdown"
}
```

Tool results skip nulls and empty lists to save tokens.

---

Previous: [Linktree page](https://lurkapi.com/docs/linkinbio/linktree.md) · Next: [Fetch a page (plain request)](https://lurkapi.com/docs/web/fetch.md)
