Skip to main content
POST
Web Scrape

Authorizations

Authorization
string
header
default:YOUR_ANYAPI_KEY
required

Your AnyAPI key as a Bearer token.

Headers

Idempotency-Key
string

Optional wallet idempotency key, scoped to this customer for 24 hours. When the gateway honors the key, this synchronous in-process execution can continue after the caller disconnects, bounded by its execution deadline. A completed replayable result charges normally exactly once and can be replayed without another provider run or charge. A pending duplicate returns 409 idempotency_in_progress; reuse with different request semantics returns 409 idempotency_conflict.

Required string length: 1 - 255

Query Parameters

fields
string

Optional. Comma-separated keys (dotted paths like author.name descend into nested objects) to keep on each result item: each row of the API's result list, or the data object itself for an API that returns one record. Keys are matched relative to that item, not against the top-level response envelope, so use jq to reshape the whole envelope. Shrinks the response without changing cost.

max_items
integer

Optional. Cap the number of result rows returned; a _truncated note reports how many were withheld so you can page via the API's own limit. An API that returns one record is not trimmed. Does not change cost.

Required range: x >= 0
summary
boolean

Optional. Return only a structural outline (top-level keys, item counts, and per-field byte sizes) instead of the full data. Does not change cost.

jq
string

Optional. A jq expression applied to the result envelope; its output replaces output (multiple outputs collect into an array). Reshape freely, e.g. jq=.data | {title, description, md: .markdown[:3500]}. Runs sandboxed with a 250ms / 2MB budget; on failure the full result is returned with a jqError. Does not change cost.

max_cost_usd
string

Optional. The most you are willing to pay for this one request, in US dollars (for example 0.05). Any route that would charge more than this is not used, so a request only ever runs on something you can afford. If nothing is available at or below your amount, the request is refused before it runs, nothing is charged, and the message tells you the cheapest price per request so you can raise it. Leave it out to accept the normal price.

Body

application/json
url
string<uri>
required

The URL of the page to scrape.

Minimum string length: 1
allowFallbacks
boolean
default:true

Optional, default true. When false, only the sources listed in source may serve; the request is refused with no charge if none of them can. When true, the listed sources are tried first and any other source may serve after them, at the normal price.

blockAds
boolean

When true (upstream default), strip ad and cookie-consent elements before capture. Set false to keep them.

excludeTags
string[]

CSS selectors to drop before capture (for example ["nav", "footer", ".ads"]). Applied after includeTags.

formats
enum<string>[]

Which representations of the page to return. Any combination of: markdown (page content as Markdown), html (the page HTML exactly as the browser received it, including head and script tags). Each requested format is returned under the matching output field. Defaults to both. rawHtml is a deprecated alias of html, returned under a rawHtml field for callers that predate the rename; send html instead.

Available options:
markdown,
html,
rawHtml
ignoreSources
string[]

Optional. Source ids to skip for this request, taken from this endpoint's lanes[].source.id. The cheapest remaining source serves and the price is that of the dearest remaining source. An id that does not serve this endpoint, or that is also in source, is rejected as invalid input with no charge; skipping every source is rejected the same way.

Minimum string length: 1
includeTags
string[]

CSS selectors to keep. When set, only content matching these selectors is captured (for example ["article", "main"] or ["#content"]).

mobile
boolean

When true, render the page with a mobile viewport and user agent instead of desktop. Some sites serve materially different content to mobile.

onlyMainContent
boolean
default:false

When true, return only the main article content, stripping navigation, headers, footers, and other boilerplate. Defaults to false to capture the full page.

preferLatencyUnderMs
integer

Optional; omit it and routing is unchanged, with the cheapest source serving. Prefer sources whose typical response time (median over the trailing 30 days, as published on this endpoint's lane health) is under this many milliseconds; among those, the cheapest serves. This can raise your price: when the cheapest source misses the target, a faster and dearer one serves, and you are quoted and charged its price. If no source is that fast the request is still served, by whichever source offers the best speed for its price - it is never refused for being slow. Sources we have not timed are tried last. This is a preference, not a guarantee: the median describes past requests and is not a ceiling on this one, and it excludes any wait this request itself asks for. On a paginated walk it applies to the first page only: later pages stay with the source that page chose, at the price it was quoted.

Required range: x >= 1
source
string[]

Optional. Source ids to prefer, in order, taken from this endpoint's lanes[].source.id in /catalog or /apis. Omit it and the cheapest source serves, with automatic failover. Listed sources are tried first in the order given, then the others, unless allowFallbacks is false. A single source with allowFallbacks false is served only by that source at its price, quoted and charged exactly, with no failover. The price is that of the dearest source that may serve. An id that does not serve this endpoint is rejected as invalid input with no charge; a listed source that is not serving right now is refused with no charge, so omit source to be served by another. On a paginated walk, later pages must include the source that served page one, or omit source.

Minimum string length: 1
waitFor
integer

Milliseconds to wait for the page to finish rendering before capture. Use this for JavaScript-heavy pages or single-page apps whose content loads after the initial paint. Capped at 15000 to stay within the request timeout. This wait is time you asked us to spend, so your response takes this much longer, and it is excluded from the latency published for this endpoint.

Required range: 0 <= x <= 15000

Response

Normalized result.

costUsd
number
required

USD charged on the original run. On a replay this value is echoed for parity; the replay itself is free.

items
integer
required

Number of result rows returned; an API that returns one record counts it as one row. For per-result APIs the per-item cost is charged against this count; for input-priced APIs the charge is per submitted input, independent of this count.

output
Web scrape output · object | null
required

Normalized output, or null when the replay payload was not retained.

provider
string
required

Always "AnyAPI".

replayed
boolean
required

True when this response replays the durable result of an earlier run without billing or upstream execution.

hint
string

Optional one-line nudge, absent when there is nothing to say. large_result: suggests the fields/max_items/summary/jq controls for a big response. paging_unavailable: means this result came from a source that cannot return a nextCursor, so it may be INCOMPLETE and cannot be continued - re-run with requireFields: ["nextCursor"] to be served only by a source that can page, which may cost more per request. malformed_query: means X read your search query differently than it looks (adjacent bare words are ANDed and OR binds tighter, so a multi-word alternative is not a phrase, and an operator after an unbracketed OR list applies to its last alternative only); the run was served and billed as X read it, and the hint carries the corrected query to send instead. date_adjusted: means a date you sent was outside the range the platform accepts, so the run used the nearest accepted date the hint names.

jqError
string

Present only when a jq expression failed; output then carries the full unshaped result and this explains why the reshape did not apply.

resultId
string

Opaque handle to the full unshaped result, cached ~15 min. Re-shape it for free (fields/max_items/summary/jq) via GET /v1/results/{id}, no re-billing. Absent when the result was too large to cache.

source
object

The source that served this run. Pass its id back as source to prefer it again, or as ignoreSources to avoid it. Absent when no source identity applies to the run.