Skip to main content
POST
Extract structured data
10 credits See the guide for examples and usage.

Authorizations

Authorization
string
header
required

Send Authorization: Bearer <API_KEY>. Keys have full access unless restricted to scopes.

Body

application/json
url
string<uri>
required

The starting website URL to crawl and extract from. Must include http:// or https://.

schema
object
required

JSON Schema (not an example object) for the returned object. Image fields include relevant page image references.

Example:
instructions
string

Optional extraction guidance, such as which facts to prioritize or how to interpret fields in the schema.

Maximum string length: 2000
factCheck
boolean
default:false

Require facts stated on the page. Unsupported fields become null or empty.

followSubdomains
boolean
default:false

When true, follow links on subdomains of the starting URL's domain.

maxPages
integer
default:5

Maximum number of pages to analyze for extraction. Hard cap: 50. Defaults to 5.

Required range: 1 <= x <= 50
maxDepth
integer

Optional maximum link depth from the starting URL (0 = only the starting page). If omitted, there is no crawl depth limit.

Required range: x >= 0
pdf
object

PDF parsing options, including whether to parse PDFs and the inclusive page range.

includeFrames
boolean
default:false

When true, iframe contents are included in Markdown before extraction.

maxAgeMs
integer
default:604800000

Return cached scrape results if a prior scrape for the same parameters is younger than this many milliseconds. Defaults to 7 days (604800000 ms).

Required range: 0 <= x <= 2592000000
waitForMs
integer

Optional browser wait time in milliseconds after initial page load for each crawled page.

Required range: 0 <= x <= 30000
settleAnimations
boolean
default:false

Wait briefly for CSS animations and transitions to settle before reading each page.

actions
(Wait · object | Perform · object | Scroll · object)[]

Browser steps before page discovery. Requires a paid plan; defaults the crawl budget to 110000 ms.

Maximum array length: 5

Browser action discriminated by do. Each variant exposes only its applicable fields.

stopAfterMs
integer
default:80000

Soft time budget for the crawl in milliseconds. Min: 10000 (10s). Max: 110000 (110s). Defaults to 80000 (80s), or 110000 (110s) when browser actions are provided.

Required range: 10000 <= x <= 110000
timeoutOpts
object

Request deadline and what to return when it passes.

zdr
enum<string>
default:disabled

enabled turns on zero data retention. Returns 403 ZDR_NOT_ENABLED unless your organization has ZDR.

Available options:
enabled,
disabled
tags
string[]

Labels for filtering usage in the dashboard.

Maximum array length: 20
Required string length: 1 - 50
Example:

Response

Successful response

status
string
required

Always ok on success.

url
string
required

The starting URL that was analyzed

urls_analyzed
string[]
required

List of URLs whose Markdown was used for extraction

data
object
required

Extracted data matching the request schema

metadata
object
required
request_id
string<uuid>
required

Unique ID of this request, also in X-Request-Id. Include it when contacting support.

Example:

"3f1c2a6e-8b4d-4c1e-9f0a-2d7b5e6c8a91"

cache_metadata
object
required

Whether this response came from cache.

partial
boolean

True when the timeout ended processing and this response contains only usable results completed so far. Unfinished results are omitted.

key_metadata
object

Credits this request used and your remaining balance.