# Scraping any web page

> Any public URL as HTML, markdown or text, past bot walls and captchas. Let Auto find the cheapest way through, from 1 credit, or pick the way and its price yourself.

- **Web page:** https://lurkapi.com/docs/web-scraping

## What it does

[`GET /v1/web/scrape`](https://lurkapi.com/docs/web/scrape.md) fetches the page at `url` and returns it in `content`, in the `format` you ask for: `html` (the default), `markdown` or `text`. Use `markdown` for LLMs: headings, links, lists and tables survive; scripts and styles don't. JSON, XML and plain-text pages come back as they are; binary files (images, PDFs) come back with `content: null`.

When a site walls the request (a JavaScript challenge, an IP block, a captcha), the call escalates, cheapest way first, until one gets through. You pay only for the one that delivered the page. Most sites need none of it: 1 credit, 1–3 seconds.

`status` is the site's own HTTP status and `url` the page's URL after redirects. In Claude and other MCP clients this is the `web_scrape` tool. [Every parameter and field](https://lurkapi.com/docs/web/scrape.md).

## Auto, a fixed price, or your own steps

Seven endpoints share one engine and one response. Auto picks the way for you; the fixed-price ones each run one way; Custom runs the ways you list.

| Endpoint | Runs | Credits |
| --- | --- | --- |
| [Auto](https://lurkapi.com/docs/web/scrape.md) `/v1/web/scrape` | the whole ladder below, cheapest first | 1–24, set by the step that got the page |
| [Fetch a page (plain request)](https://lurkapi.com/docs/web/fetch.md) `/v1/web/fetch` | `direct` then `static` | 1 credit |
| [Render a page (browser)](https://lurkapi.com/docs/web/render.md) `/v1/web/render` | `browser` | 4 credits; 14 if a captcha is solved |
| [Unblock a page (partner browser)](https://lurkapi.com/docs/web/unblock.md) `/v1/web/unblock` | `unblocker` | 4 credits |
| [Fetch a page from a residential IP](https://lurkapi.com/docs/web/residential.md) `/v1/web/residential` | `residential` | 5 credits |
| [Render a page from a residential IP](https://lurkapi.com/docs/web/render-residential.md) `/v1/web/render-residential` | `browser_residential` | 14 credits; 24 if a captcha is solved |
| [Custom](https://lurkapi.com/docs/web/custom.md) `/v1/web/custom` | the steps you list in `steps`, in your order | set by the step that got the page, at its fixed price |

Every answer says what it cost in `billing`: `mode` (`auto`, `fixed` or `custom`), `credits` (the same as `credits_charged`), a `breakdown` (`base`, `render`, `premium`, `premium_render`, `captcha`) and, on Auto and Custom, a `note` on why the price varies. The fixed-price endpoints and Custom solve a captcha (+10) only with `captcha=true`, so their price stays what you chose unless you allow one.

Custom is for when you know a site: `steps=fetch,unblock` never costs more than 4 credits and never runs JavaScript; `steps=render,render-residential` always returns the page as rendered. It skips a step a wall has already ruled out, the way Auto does.

## Try it live

Pick an endpoint and run it: real data, no signup, 5 free requests a day.

Live playground: https://lurkapi.com/docs/web#try

## How it gets the page

An Auto call climbs this ladder and stops at the first step that gets the page:

| Step | How | For | Extra credits |
| --- | --- | --- | --- |
| `direct` | A plain request | Most sites | none |
| `static` | A static ISP IP with a real Chrome fingerprint (TLS, HTTP/2 and headers) | Sites that refuse a datacenter IP or a client that isn't a browser | none |
| `browser` | A real Chrome on a static ISP IP | JavaScript challenges (Cloudflare, AWS WAF, Akamai, PerimeterX) and pages that are only a JavaScript app | +3 |
| `unblocker` | A partner's browser on its own IPs | Sites that refuse our IPs (DataDome and the like). It returns the HTML the site sent, before JavaScript runs, and the URL you asked for, not where redirects ended. | +3 |
| `residential` | A residential IP | The last resort, when the steps above are refused | +4 |
| `browser_residential` | Our browser on a residential IP | The last resort for pages that need a browser | +13 |
| captcha | A captcha solver, inside our browser | A Cloudflare Turnstile, DataDome slider or reCAPTCHA our browser can't click past | +10, only when one is solved |

What stopped the last step decides the next: a JavaScript challenge skips the remaining plain requests for a browser, an IP block (DataDome, a bare 403 or 429) skips our IPs for the partner's browser, and an empty JavaScript app goes to our browser. So a Cloudflare-challenged site typically goes `direct → static → browser`, and a DataDome site `direct → static → unblocker`.

`via` names the step that delivered the page, and `captcha` the captcha solved for it (`turnstile`, `datadome` or `recaptcha`), or `null`. `attempts` lists every step tried, in order: how long it took (`ms`), the HTTP `status` it got, and what walled it (`wall`, like `cloudflare challenge` or `datadome block`) or why it failed (`error`).

A DataDome site's first call: Yelp, in production (trimmed):

```json
{
  "credits_charged": 4,
  "billing": { "mode": "auto", "credits": 4, "breakdown": { "base": 1, "render": 3 }, "note": "Auto: the price follows the strategy that got the page…" },
  "via": "unblocker",
  "captcha": null,
  "attempts": [
    { "via": "direct", "ms": 43, "status": 403, "wall": "datadome block", "error": null },
    { "via": "static", "ms": 802, "status": 403, "wall": "datadome block", "error": null },
    { "via": "unblocker", "ms": 17651, "status": 200, "wall": null, "error": null }
  ]
}
```

## Credits

- **1 credit for every page delivered.** Paid steps add to it: **+3** when a browser renders it (ours or the partner's), **+4** for a residential IP, **+13** for our browser on one, **+10** when a captcha is solved. A call costs 1–24 credits.
- **You pay for the step that got the page.** A paid step is held from your balance before it runs and handed back if it doesn't deliver, so a render that hit a wall followed by a residential IP that got the page costs 5, not 8.
- **Failures are free:** walled at every step, the site down, the site's own `5xx`, or no answer in time.
- **A page is a page:** the site's own 404 or 410 comes back with that `status` and is charged.
- **Paid steps need the credits when they start.** If your balance can't cover a step the site needs, the call stops with `402 insufficient_credits`, free. A captcha solve is the exception: our browser still renders, without the solver.

| The site | Usually delivered by | Credits |
| --- | --- | --- |
| Most sites | `direct` or `static` | 1 |
| A JavaScript challenge, or a page that's only a JavaScript app | `browser` | 4 |
| Refuses our IPs (DataDome and the like) | `unblocker` | 4 |
| A captcha our browser can't click past | `browser` and a solve | 14 |
| Only a residential IP gets through | `residential` or `browser_residential` | 5 or 14 |
| Only a residential IP gets through, behind a captcha | `browser_residential` and a solve | 24 |

## Controlling cost and behaviour

On Auto, the defaults escalate as far as the site makes them. These parameters cap or force the ladder (or use a fixed-price endpoint, or Custom):

| Parameter | Values | What it does |
| --- | --- | --- |
| `render` | `auto` (default), `true`, `false` | `auto`: a browser only when the page needs one. `true`: always our browser, so `content` is the page after JavaScript ran (never the partner's browser, which returns the HTML as sent). `false`: plain requests only. |
| `proxy` | `auto` (default), `static`, `residential` | `auto`: static ISP IPs and the partner's browser; a residential IP only when they're refused. `static`: never residential. `residential`: residential IPs only. |
| `country` | A two-letter code: `us`, `gb`, `de`… | Fetch from a residential IP in that country. Implies `proxy=residential`. |
| `captcha` | `true` (default), `false` | `false`: never pay for a solve. |
| `wait_for` | A CSS selector, like `.price` | Wait until it's on the page before returning. Implies `render=true`. |
| `wait` | Milliseconds, up to 10000 | Let the page settle this long after it loads. Implies `render=true`. |

What each choice can cost, in credits:

| Parameters | Steps allowed | Least | Most |
| --- | --- | --- | --- |
| The defaults | all | 1 | 24 |
| `captcha=false` | all, without solves | 1 | 14 |
| `proxy=static` | `direct`, `static`, `browser`, `unblocker` | 1 | 14 |
| `proxy=static&captcha=false` | the same, without solves | 1 | 4 |
| `render=false` | `direct`, `static`, `residential` | 1 | 5 |
| `render=false&proxy=static` | `direct`, `static` | 1 | 1 |
| `render=true`, `wait_for` or `wait` | `browser`, `browser_residential` | 4 | 24 |
| `proxy=residential` or `country` | `residential`, `browser_residential` | 5 | 24 |

So `render=false&proxy=static` never costs more than 1 credit, and `render=true` always returns the page after JavaScript ran. A narrower ladder can fail where the defaults would get through; after `render=false`, `proxy=static` or `captcha=false`, the error says which to relax.

## Speed

- **Most pages:** 1–3 s.
- **Rendered by our browser:** 5–15 s.
- **Through the partner's browser:** 15–35 s.
- **Up to 85 s** for a solved captcha or the slowest sites. Give your HTTP client a timeout of at least 90 s; the API answers `504 upstream_timeout` by then.

Repeat calls get faster:

- A site that needed a stronger step starts there for the next 15 minutes (for calls allowed the same steps), instead of hitting the same walls again.
- Once our browser gets past a site's wall, later calls to that site reuse the pass for up to 30 minutes: often as a plain request (1 credit, a few seconds), else as a render that skips the wall, so a captcha isn't solved, or paid for, twice.

## Failures

A call that gets no page fails with `502 upstream_error` and costs nothing. The `error` says why:

- **Walled at every step it was allowed.** The message names each step and its wall, and the options that might get through (`Allowing render=auto and proxy=auto may get through.`). The same call (same `url` and options) then fails at once for 5 minutes; change `render` or `proxy` to try other steps.
- **The site is down:** it doesn't resolve, its server is down behind Cloudflare (`521`–`530`), it refuses connections, its certificate is invalid, or no network got an answer. Every call to that host then fails at once for 3 minutes.
- **The site answered `5xx`**, or **didn't answer in time**: retry.

| Status | Code | When |
| --- | --- | --- |
| 400 | `invalid_params` | A missing or invalid parameter, or a `url` we won't fetch: not http(s), a port other than 80 or 443, credentials in it, or a host that isn't public (`localhost`, a private IP, or a name that resolves to one). `issues` names it. |
| 502 | `upstream_error` | No page: walled at every step allowed, the site down, its `5xx`, or no answer. Free. |
| 504 | `upstream_timeout` | The call ran past 90 s. Free; rare, since the ladder stops at 85 s. |

## Limits and good use

- **Public pages only.** The endpoint takes no cookies or headers from you, so nothing behind a login.
- **Public hosts on ports 80 and 443**, in URLs up to 2,048 characters.
- **Encode `url`** like any query value: `curl -G --data-urlencode "url=…"` does it for you.
- **Pages over 2 MB are cut** there, and `truncated` is `true`.
- **Not cached.** Every call fetches the page and costs credits; keep your own copy if you need a page twice.
- **Use `format=markdown` for LLMs.** It's much smaller than the HTML and keeps the structure a model needs.

## Examples

### A page as markdown

curl:

```bash
curl "https://api.lurkapi.com/v1/web/scrape?url=https%3A%2F%2Fbooks.toscrape.com%2F&format=markdown" \
  -H "x-api-key: YOUR_API_KEY"
```

JavaScript:

```js
const params = new URLSearchParams({
  url: "https://books.toscrape.com/",
  format: "markdown",
});
const res = await fetch(`https://api.lurkapi.com/v1/web/scrape?${params}`, {
  headers: { "x-api-key": process.env.LURKAPI_KEY },
});
const data = await res.json();
if (!data.success) throw new Error(`${data.code}: ${data.error}`);
console.log(data.content);
```

Python:

```python
import os
import requests

res = requests.get(
    "https://api.lurkapi.com/v1/web/scrape",
    params={
        "url": "https://books.toscrape.com/",
        "format": "markdown",
    },
    headers={"x-api-key": os.environ["LURKAPI_KEY"]},
    timeout=95,
)
data = res.json()
if not data["success"]:
    raise RuntimeError(f"{data['code']}: {data['error']}")
print(data["content"])
```

### A walled site

Nothing to set: the defaults escalate until the page comes through. A URL with its own query string has to be encoded:

curl:

```bash
curl -G "https://api.lurkapi.com/v1/web/scrape" \
  --data-urlencode "url=https://www.tripadvisor.com/Search?q=lisbon" \
  -d format=markdown \
  -H "x-api-key: YOUR_API_KEY"
```

### Always 1 credit

Plain requests only: a walled site fails, free, instead of escalating.

curl:

```bash
curl -G "https://api.lurkapi.com/v1/web/fetch" \
  --data-urlencode "url=https://news.ycombinator.com/" \
  -d format=text \
  -H "x-api-key: YOUR_API_KEY"
```

### Your own steps, at most 4 credits

A plain request first, then the partner's browser if the site walls it:

curl:

```bash
curl -G "https://api.lurkapi.com/v1/web/custom" \
  --data-urlencode "url=https://www.yelp.com/biz/the-house-san-francisco" \
  -d steps=fetch,unblock \
  -d format=markdown \
  -H "x-api-key: YOUR_API_KEY"
```

[Get a free API key](https://lurkapi.com/login?next=/dashboard)

## FAQ

### Why did I pay 14 credits?

Our browser rendered the page (+3) and a captcha solver cleared a captcha it couldn't click past (+10); `captcha` in the response names it. Pass `captcha=false` to never pay for a solve: the call then tries the other steps and fails, free, if none gets through.

### Why is a site “down”?

The site itself isn't answering: it doesn't resolve, its server is down behind Cloudflare, it refuses connections, its certificate is invalid, or no network got an answer. No other IP or browser changes that, so calls to it fail at once, free, for 3 minutes; then we try it again.

### Can I scrape logged-in pages?

No. It fetches public pages only and takes no cookies or headers from you. A page behind a login comes back as whatever the site shows a visitor who isn't signed in, usually its login page.

### How is this different from the platform endpoints?

The [platform endpoints](https://lurkapi.com/docs.md#endpoints) return structured JSON fields at a listed price per call, and are cached. Web scrape returns any page as HTML, markdown or text, fresh every time, and costs what the site made it take (1–24 credits). If a platform endpoint covers what you need, use it.
