API reference
Document Parser API
Extract text and metadata from HTML, plain text, Markdown, CSV/TSV, JSON, XML and PDF. Read visible text from one static JPEG, PNG or WebP image up to 2 MB and 16 megapixels, including TikTok/Instagram post covers via data.cover_ocr. Empty PDF extraction is uncharged. Office files, archives and unsupported image formats return 422.
Endpoint
/v1/web/parselive · proven2 creditsParameters
Query parameters
| Name | Required | Description | Example |
|---|---|---|---|
url | yes | Public document URL. Also accepts one static JPEG, PNG or WebP image up to 2 MB and 16 megapixels for text extraction; other slides or video frames require separate requests. | https://www.w3.org/WAI/ER/tests/xhtml/testfiles/resources/pdf/dummy.pdf |
timeout | no | Fetch budget in milliseconds, clamped to 1000-15000 | 10000 |
dry_run | no | Read a zero-credit estimate without fetching sources, running AI, or reserving credits. Cache status is a snapshot, not a guarantee at execution. | 1 |
Examples
Make the request
curl "https://www.monocrawl.com/v1/web/parse?url=https%3A%2F%2Fwww.w3.org%2FWAI%2FER%2Ftests%2Fxhtml%2Ftestfiles%2Fresources%2Fpdf%2Fdummy.pdf" \ -H "x-api-key: mn_your_key_here"
TypeScript
const key = process.env.MONOCRAWL_API_KEY;
if (!key) throw new Error('Set MONOCRAWL_API_KEY on your server.');
const res = await fetch(
"https://www.monocrawl.com/v1/web/parse?url=https%3A%2F%2Fwww.w3.org%2FWAI%2FER%2Ftests%2Fxhtml%2Ftestfiles%2Fresources%2Fpdf%2Fdummy.pdf",
{ headers: { "x-api-key": key } },
);
const body = await res.json();
if (!body.success) {
// one error shape for every endpoint — see /docs/errors
throw new Error(`${body.error.type}: ${body.error.message}`);
}
console.log(body.data, "credits left:", body.credits_remaining);Python
import os
import requests
res = requests.get(
"https://www.monocrawl.com/v1/web/parse?url=https%3A%2F%2Fwww.w3.org%2FWAI%2FER%2Ftests%2Fxhtml%2Ftestfiles%2Fresources%2Fpdf%2Fdummy.pdf",
headers={"x-api-key": os.environ["MONOCRAWL_API_KEY"]},
timeout=60,
)
body = res.json()
if not body["success"]:
# one error shape for every endpoint — see /docs/errors
raise RuntimeError(f"{body['error']['type']}: {body['error']['message']}")
print(body["data"], "credits left:", body["credits_remaining"])Response
Response fields and example
This example is illustrative, not a captured live response. Variable-cost operations may settle a charge different from the list price below. A successful response puts the platform payload in data and reports the exact credits used, remaining balance, request id and cache status beside it.
{
"success": true,
"platform": "web",
"endpoint": "/v1/web/parse",
"data": {
"detected_type": "image",
"text": "Example visible text",
"texts": [
"Example visible text"
],
"truncated": false
},
"credits_used": 2,
"credits_remaining": 99,
"request_id": "req_…",
"cached": false
}The example is illustrative and shows a documented subset of data. The fields below are optional across supported sources; nullable fields can also be absent. Preserve unknown values and accept additional fields.
Schema basis: reviewed response shaper. See schema coverage and validation limits.
Download the data JSON Schema. Validate response.data, not the whole envelope. A valid shape does not establish that every field or source record was returned.
| Field inside data | Type | Meaning |
|---|---|---|
url | string | null | url |
final_url | string | null | final url |
content_type | string | null | content type |
detected_type | string | null | detected type |
detected_via | string | null | detected via |
text | string | null | text |
byte_size | number | null | byte size; null is unknown. |
text_length | number | null | text length; null is unknown. |
pages | number | null | pages; null is unknown. |
metadata | object | null | Nested response fields; optional unless explicitly documented. |
truncated | boolean | null | Whether returned document text was cut short. |
texts | array | Nested response fields; optional unless explicitly documented. |
_warnings | array | Nested response fields; optional unless explicitly documented. |
Static JPEG, PNG and WebP: one image up to 2 MB and 16 megapixels, normalized to fit 1024 pixels. Other carousel slides and video frames require separate requests. OCR can be inaccurate; output truncation fails rather than silently returning an incomplete transcription.
HTML, text and PDF parsing remain available without a model call. The existing endpoint price applies.
Failures use the typed error envelope. A confirmed uncharged or refunded failure reports zero; a pending reconciliation can report an unknown charge. Read credits_used and error.details.billing_status, and keep request_id for recovery. Response contract · Error reference