Skip to main content
GET
Scrape Markdown
1 Credit With actions: 2 Credits

Authorizations

Authorization
string
header
required

Bearer authentication header of the form Bearer <API_KEY>, where <API_KEY> is your api key.

Query Parameters

url
string<uri>
required

Full URL to scrape into LLM usable Markdown (must include http:// or https:// protocol)

Minimum string length: 1

Preserve hyperlinks in Markdown output

includeImages
boolean
default:false

Include image references in Markdown output

shortenBase64Images
boolean
default:true

Shorten base64-encoded image data in the Markdown output

useMainContentOnly
boolean
default:false

Extract only the main content of the page, excluding headers, footers, sidebars, and navigation

includeHTML
boolean
default:false

When true, the response also includes an html field with the page HTML the Markdown was converted from — the same body the Scrape HTML endpoint returns for the equivalent request.

pdf
object

PDF parsing controls. Use start/end to limit text extraction and embedded-image detection/OCR to an inclusive 1-based page range.

includeFrames
boolean
default:false

When true, the contents of iframes are rendered to Markdown.

includeSelectors
string[] | null

CSS selectors. When provided, only matching HTML subtrees (and their descendants) are kept before conversion to Markdown. When omitted, the entire document is kept. Examples: "article.main", "#content", "[role=main]".

Maximum array length: 50
Required string length: 1 - 2048
excludeSelectors
string[] | null

CSS selectors to remove before conversion to Markdown. Applied after includeSelectors. Exclusion takes precedence: an element matching both is removed. Examples: "nav", "footer", ".ad-banner", "[aria-hidden=true]".

Maximum array length: 50
Required string length: 1 - 2048
maxAgeMs
integer | null
default:86400000

Return a cached result if a prior scrape for the same parameters exists and is younger than this many milliseconds. Defaults to 1 day (86400000 ms) when omitted. Max is 30 days (2592000000 ms). Set to 0 to always scrape fresh.

Required range: 0 <= x <= 2592000000
waitForMs
integer | null

Optional browser wait time in milliseconds after initial page load before converting the page to Markdown. Min: 0. Max: 30000 (30 seconds).

Required range: 0 <= x <= 30000
settleAnimations
boolean
default:false

When true, waits briefly for CSS and transition animations to settle before converting to Markdown. Defaults to false. This adds a bit of latency in exchange for more stable output on animated pages.

actions
(Wait · object | Perform · object | Scroll · object)[] | null

Optional browser actions executed in array order after the page loads and before content is captured. Requires a paid plan. Send a JSON array in the query parameter. Maximum: 5 actions.

Maximum array length: 5

Browser action discriminated by do. Each variant exposes only its applicable fields.

headers
object

Optional outbound HTTP headers forwarded only to the target URL, sent as deep-object query params such as headers[X-Custom]=value. When provided, caching is bypassed: the result is neither read from nor written to cache.

country
enum<string>

Fetch the target page through a residential proxy in this country (ISO 3166-1 alpha-2).

Available options:
ad,
ae,
af,
ag,
ai,
al,
am,
ao,
ar,
at,
au,
aw,
az,
ba,
bb,
bd,
be,
bf,
bg,
bh,
bi,
bj,
bm,
bn,
bo,
bq,
br,
bs,
bw,
by,
bz,
ca,
cd,
cf,
cg,
ch,
ci,
cl,
cm,
cn,
co,
cr,
cv,
cw,
cy,
cz,
de,
dj,
dk,
dm,
do,
dz,
ec,
ee,
eg,
es,
et,
fi,
fj,
fr,
ga,
gb,
gd,
ge,
gf,
gg,
gh,
gm,
gn,
gp,
gq,
gr,
gt,
gu,
gw,
gy,
hk,
hn,
hr,
ht,
hu,
id,
ie,
il,
im,
in,
iq,
ir,
is,
it,
je,
jm,
jo,
jp,
ke,
kg,
kh,
kn,
kr,
kw,
ky,
kz,
la,
lb,
lc,
lk,
lr,
ls,
lt,
lu,
lv,
ly,
ma,
mc,
md,
me,
mf,
mg,
mk,
ml,
mm,
mn,
mo,
mq,
mr,
mt,
mu,
mv,
mw,
mx,
my,
mz,
na,
nc,
ne,
ng,
ni,
nl,
no,
np,
nz,
om,
pa,
pe,
pf,
pg,
ph,
pk,
pl,
pr,
ps,
pt,
py,
qa,
re,
ro,
rs,
ru,
rw,
sa,
sc,
sd,
se,
sg,
si,
sk,
sl,
sm,
sn,
so,
sr,
ss,
st,
sv,
sx,
sy,
sz,
tc,
td,
tg,
th,
tj,
tl,
tm,
tn,
tr,
tt,
tw,
tz,
ua,
ug,
us,
uy,
uz,
vc,
ve,
vg,
vi,
vn,
ye,
yt,
za,
zm,
zw
Example:

"de"

timeoutMS
integer

Optional timeout in milliseconds for the request. If the request takes longer than this value, it will be aborted with a 408 status code. Maximum allowed value is 300000ms (5 minutes).

Required range: 1 <= x <= 300000
zdr
enum<string>
default:disabled

Set to enabled to bypass shared caches and omit request and response content from retained usage logs. Requires zero data retention to be enabled for your organization (contact [email protected]), otherwise the request fails with ZDR_NOT_ENABLED. Successful ZDR responses include X-Context-ZDR: true.

Available options:
enabled,
disabled
tags
string[]

Comma-separated tags for tracking request usage. Up to 20 tags, each 1-50 characters. Optional tags for tracking usage. Up to 20 tags, each 1 to 50 characters.

Maximum array length: 20
Required string length: 1 - 50
Example:

Response

Successful response

success
enum<boolean>
required

Indicates success

Available options:
true
markdown
string
required

Page content converted to GitHub Flavored Markdown

contentLength
integer
required

UTF-8 byte length of the returned Markdown. Use 0 to identify an empty result and compare small values against your workload's minimum useful-content threshold.

Required range: x >= 0
url
string
required

The URL that was scraped

metadata
object
required

Metadata extracted from the scraped page HTML.

cache_metadata
object
required

Cache outcome for this response. Composite responses are hits only when every cache-controlled fetch contributing to the output was a hit; age_ms is the oldest contributing hit.

html
string

Only present when includeHTML=true: the page HTML the Markdown was converted from — the same body the Scrape HTML endpoint returns for the equivalent request.

key_metadata
object

Credit usage, included whenever a valid API key is provided.

actionsApplied
object[]

One verified outcome per requested browser action, in request order.

actionsHtmlStale
boolean

True when an action was applied but the returned content could not be refreshed afterward.