## web/crawl

`GET /v1/web/crawl`

Start a breadth-first same-site crawl and get a job id back immediately; poll /v1/web/jobs/{job_id} for the collected pages (url, title, markdown excerpt, status).

*Start an async web crawl*

| Parameter | Required | Description |
| --- | --- | --- |
| `url` | yes | Seed URL to crawl (http/https, public hosts only) — e.g. `https://example.com` |
| `limit` | no | Maximum pages to fetch (default 10, cap 50) — e.g. `10` |
| `max_depth` | no | Link hops from the seed page (default 2, cap 3) — e.g. `2` |
| `include_paths` | no | Comma-separated glob patterns; a discovered path must match at least one. * matches any characters, ^ and $ anchor, everything else is literal — e.g. `^/blog/.*` |
| `exclude_paths` | no | Comma-separated glob patterns; a discovered path matching any of them is skipped — e.g. `/tag/*,/author/*` |
| `formats` | no | Comma-separated subset of markdown,text,links (default markdown) — e.g. `markdown` |
| `only_main_content` | no | Render only the main content region of each page (default true) — e.g. `true` |
| `dry_run` | no | Read a zero-credit estimate without fetching sources, running AI, or reserving credits. Cache status is a snapshot, not a guarantee at execution. — e.g. `1` |

**Cost:** 5 credits at list price per successful uncached response; variable-cost operations may quote or settle a different charge. Confirmed uncharged or refunded failures report zero. Pending reconciliation can report credits_used:null; retain request_id and the original idempotency key.

```bash
curl "https://www.monocrawl.com/v1/web/crawl?url=https%3A%2F%2Fexample.com" \
  -H "x-api-key: mn_your_key_here"
```

### Response data

Illustrative abbreviated data; documented core fields are optional and may be null. Additional source fields are allowed.

```json
{
  "id": "job_example",
  "status": "queued"
}
```

[Field reference](https://www.monocrawl.com/docs/endpoints/web/crawl#response) · [JSON Schema for response.data](https://www.monocrawl.com/schemas/web/crawl.json)
