Scraping API
https://scraping-api.datafuel.ai/api/v1
Help centerOpenAPI spec

Jobs

Asynchronous multi-target batches

GET/job

List your jobs and crawls

Your jobs newest first, crawls included (type: crawl), with the same counters as GET /job/{job_id}. Pass next_cursor back as cursor for the next page; it is absent on the last page. Unknown status or type values match nothing.

Query
status
string
pendingprocessingcompletedcompleted_with_errorsfailedcancelled
type
string
unlockerllm_scrapingserpmapcrawl
start_date
string
YYYY-MM-DD (UTC), created on or after this day. Wrong format answers 400 INVALID_DATE_FORMAT.
end_date
string
YYYY-MM-DD (UTC), created on or before this day (inclusive). Before start_date answers 400 INVALID_DATE_RANGE.
limit
integer
Items per page; values above 200 are capped. Negative or not a number answers 400 INVALID_QUERY_PARAM. Default: 50.
cursor
string
next_cursor of the previous page. A cursor that does not decode answers 400 INVALID_CURSOR.
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "jobs": [
    {
      "id": "1d2c3b4a-5e6f-4a7b-8c9d-0e1f2a3b4c5d",
      "type": "crawl",
      "status": "completed",
      "tasks_count": 120,
      "tasks_done": 120,
      "tasks_remaining": 0,
      "total_cost": 120,
      "created_at": "2026-09-29T10:00:00Z",
      "updated_at": "2026-09-29T10:04:31Z"
    }
  ]
}
{
  "code": "INVALID_QUERY_PARAM",
  "message": "limit, page and job_id must be valid values"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
POST/jobtype: unlocker · asynchronous batch

Unlocker

Queues a batch of targets and returns immediately with {"id": ...}. Attribute shape matches the task variant, except the target field is plural: urls (unlocker), prompts (llm_scraping) or queries (serp). At least two targets are required (400 JOB_REQUIRES_MULTIPLE_TARGETS); an empty one answers 400 MISSING_TARGET. map is not a job type (400 UNSUPPORTED_TASK_TYPE) and crawls go through POST /crawl. Every task is charged when the job is queued.

Headers
Idempotency-Key
string
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
proxy_city
string
proxy_state
string
proxy_asn
string
multithreaded
boolean
Process the batch concurrently (up to your concurrency limit) instead of one task at a time Default: false.
Attributes
urlsREQUIRED
array<string>
method
string
Case-insensitive; anything else answers 400 INVALID_ATTRIBUTES. Default: "GET".
GETPOSTPUTDELETEPATCHHEADOPTIONS
body
string
Raw request body sent to the target with method POST, PUT or PATCH, on both engines (js_rendering on or off).
content_type
string
Content-Type header sent with body. A Content-Type in headers takes precedence.
headers
object
header_order
array<string>
cookie_string
string
user_agent_type
string
chromefirefoxsafariedge
user_agent
string
Explicit User-Agent, overrides user_agent_type
js_rendering
boolean
Render the page in a real browser (browser pricing, slower). Off by default: a plain HTTP fetch keeps the exact URL you asked for; sites that redirect filtered URLs to a canonical page or render listings client-side need it on. Check redirected / final_url in the response either way. Default: false.
js_instructions
array<object> | object
Browser actions run after load (js_rendering only). Recommended form: an array of single-action objects, run in order, where an action may repeat, e.g. [{"fill": ["input[name=q]", "laptops"]}, {"click": "button[type=submit]"}, {"wait_ms": 1000}]. The older object form keyed by action ({"fill": [...], "click": "..."}) is still accepted, but its order is not guaranteed and an action cannot repeat. The supported actions and their argument shapes are listed by GET /config/js-instructions; actions marked iframe also accept the iframe_ prefix. Anything else (a string, a number, an array element that is not a single-action object, an unknown action) answers 400 INVALID_ATTRIBUTES.
block_resource
string | array<string>
Resource types to block in the browser: one value as a string, or several as an array. An unknown value answers 400 INVALID_ATTRIBUTES.
wait_for_selector
string
js_rendering only: CSS selector to wait for before capturing
wait_for_selector_timeout_ms
integer
js_rendering only: timeout of wait_for_selector Default: 3000.
extract_selector
string
Extract only matching elements instead of formatting the whole page. Always a string: either a bare CSS selector (result key selector) or a JSON-encoded object of name → selector for several fields. A JSON object (not a string) answers 400 INVALID_ATTRIBUTES. Append @attr to read an attribute instead of the text. The result is {"data": {name: [values...]}} and result_format is forced to json.
extract_regex
string
Regex extraction on the raw HTML. Always a string: either a bare RE2 pattern (result key regex) or a JSON-encoded object of name → pattern. The first capture group is returned when present, otherwise the whole match; all matches are collected as a list.
result_format
string
html: raw page. markdown: cleaned text with a title/meta header, best for LLMs. json: schema.org / JSON-LD and embedded JSON objects. png / jpeg: full-page screenshot, base64; needs js_rendering: true (the request engine returns html instead). Default: "html".
htmlmarkdownjsonpngjpeg
result_template
string
Comma-separated built-in extractors returning JSON: images, links, headings, phone_numbers, emails, meta_tags, tables, schema_org, all
include_images
boolean
Markdown only. false drops ![](src) images (and the links that only wrapped them); signed image URLs often dominate the token count of a listing page. The MCP server defaults this to false. Default: true.
main_content_only
boolean
Markdown only. Always render just the <main> / <article> container when the page has one. Default: false.
result_use_ai
boolean
Post-process the raw result with an LLM
result_ai_format
object
JSON shape the AI output must follow
result_ai_prompt
string
ai_provider
string
openaianthropicgoogle
ai_model
string
ai_api_key
string
Bring-your-own-key for AI post-processing; stored encrypted
proxy_session_id
string
Sticky session. Requests with the same id reuse the same proxy exit.
proxy_ttl
integer
Sticky session lifetime in seconds
proxy_type
string
Filled from the top-level proxy_type. Set the plan at the top level; that is the one that is billed.
proxy_country
string
Overrides the top-level proxy_country
proxy_city
string
Overrides the top-level proxy_city
proxy_state
string
Overrides the top-level proxy_state
proxy_asn
string
Overrides the top-level proxy_asn
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "unlocker",
    "multithreaded": true,
    "proxy_type": "Basic",
    "attributes": {
      "urls": [
        "https://example.com/product/1",
        "https://example.com/product/2"
      ],
      "method": "GET",
      "js_rendering": false
    }
  }'
Response
{
  "id": "9f1d3c2a-4b6e-4f8a-9c7d-2e5b8a1f0d43"
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
POST/jobtype: llm_scraping · asynchronous batch

LLM scraping

Queues a batch of targets and returns immediately with {"id": ...}. Attribute shape matches the task variant, except the target field is plural: urls (unlocker), prompts (llm_scraping) or queries (serp). At least two targets are required (400 JOB_REQUIRES_MULTIPLE_TARGETS); an empty one answers 400 MISSING_TARGET. map is not a job type (400 UNSUPPORTED_TASK_TYPE) and crawls go through POST /crawl. Every task is charged when the job is queued.

Headers
Idempotency-Key
string
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
proxy_city
string
proxy_state
string
proxy_asn
string
multithreaded
boolean
Process the batch concurrently (up to your concurrency limit) instead of one task at a time Default: false.
Attributes
promptsREQUIRED
array<string>
engineREQUIRED
string
Every engine the API implements. Which ones accept work right now is in GET /config/capabilities.
openaigeminigoogle_ai_modeperplexitycopilot
websearch
boolean
Let the engine browse the web while answering
follow_up_prompt
string
Second prompt sent in the same conversation
proxy_country
string
Exit country; falls back to the top-level proxy_country
location
string
Location the engine should answer for (used by gemini and perplexity); defaults to the exit country
result_format
string
Ignored; LLM scraping always returns json.
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "llm_scraping",
    "multithreaded": true,
    "attributes": {
      "prompts": [
        "best CRM tools for startups 2026",
        "top project management software 2026"
      ],
      "engine": "perplexity",
      "websearch": true
    }
  }'
Response
{
  "id": "9f1d3c2a-4b6e-4f8a-9c7d-2e5b8a1f0d43"
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
POST/jobtype: serp · asynchronous batch

SERP

Queues a batch of targets and returns immediately with {"id": ...}. Attribute shape matches the task variant, except the target field is plural: urls (unlocker), prompts (llm_scraping) or queries (serp). At least two targets are required (400 JOB_REQUIRES_MULTIPLE_TARGETS); an empty one answers 400 MISSING_TARGET. map is not a job type (400 UNSUPPORTED_TASK_TYPE) and crawls go through POST /crawl. Every task is charged when the job is queued.

Headers
Idempotency-Key
string
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
proxy_city
string
proxy_state
string
proxy_asn
string
multithreaded
boolean
Process the batch concurrently (up to your concurrency limit) instead of one task at a time Default: false.
Attributes
queriesREQUIRED
array<string>
One search per query, billed per query.
country
string
Two-letter country for Google results (gl), e.g. us, gb, fr.
language
string
Language code (hl), e.g. en, en-gb, es-419.
page
integer
1-based result page; offset is (page-1) * 10 (start). Default: 1.
google_domain
string
Google domain to query, e.g. google.co.uk (a leading www. is dropped). Default google.com; pooled per domain.
location
string
Free-text origin of the search, sent to Google as uule. Exclusive with uule and lat/lon. Google may still place the local pack by the exit IP.
uule
string
Google-encoded location. Exclusive with location, lat/lon.
lat
number
GPS latitude of the search origin, -90 to 90. Requires lon.
lon
number
GPS longitude, -180 to 180. Requires lat.
radius
integer
Bias radius in meters, carried inside the GPS location. Needs one of location, uule or lat/lon.
cr
string
Limit to countries, |-delimited upper-case codes, e.g. countryFR|countryDE.
lr
string
Limit to languages, |-delimited, e.g. lang_fr|lang_de.
tbs
string
Advanced "to be searched" filters (dates, news, patents, ...).
safe
string
SafeSearch. active filters explicit content, off disables.
activeoff
nfpr
boolean
true excludes auto-corrected-query results (served via Google's own link).
filter
boolean
Similar/omitted-results filter. true on (default); sends 0 only when explicitly false.
uds
string
Google filter string, echoed back per filter in the results.
kgmid
string
Knowledge Graph listing id.
si
string
Cached/encrypted search-params blob (Knowledge Graph tabs).
ludocid
string
Google CID of a place.
lsig
string
Forces the knowledge-graph map view to render.
ibp
string
Layout/expansion renderer, e.g. gwp;0,7.
proxy_country
string
Accepted but not used yet. SERP searches exit through DataFuel's own pool; use country and language to target results.
result_format
string
json structured results, markdown LLM-optimized, html raw served page. Default: "json".
jsonmarkdownhtml
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "serp",
    "multithreaded": true,
    "proxy_country": "US",
    "attributes": {
      "queries": [
        "best CRM tools for startups 2026",
        "top project management software 2026"
      ],
      "country": "us",
      "language": "en"
    }
  }'
Response
{
  "id": "9f1d3c2a-4b6e-4f8a-9c7d-2e5b8a1f0d43"
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
GET/job/{job_id}

Get job progress

Path
job_idREQUIRED
string
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job/{job_id} \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "status": "processing",
  "tasks_count": 50,
  "tasks_done": 32,
  "tasks_remaining": 18,
  "total_cost": 64
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "JOB_NOT_FOUND",
  "message": "Job not found"
}
GET/job/{job_id}/results

Get job results

Path
job_idREQUIRED
string
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job/{job_id}/results \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "tasks_count": 50,
  "tasks_failed": 1,
  "tasks_complete": 49,
  "tasks_result": [
    {
      "id": "0c6f1f9e-2b3a-4c5d-8e9f-1a2b3c4d5e6f",
      "status": "completed",
      "status_code": 200,
      "credits_used": 1,
      "result": {
        "data": "# Example Domain\n\nThis domain is for use in documentation examples."
      }
    },
    {
      "id": "5d4c3b2a-1f0e-4d9c-8b7a-6f5e4d3c2b1a",
      "status": "failed",
      "blocked": true,
      "status_code": 403,
      "credits_used": 0,
      "result": {
        "status": "failed",
        "error_detail": "target answered 403 (anti-bot)",
        "status_code": 403
      }
    }
  ]
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "JOB_NOT_FOUND",
  "message": "Job not found"
}
POST/job/{job_id}/cancel

Cancel a job

Stops the job. Tasks that have not started yet fail and their credits are refunded right away; tasks already running finish and are billed as usual. The job ends with status cancelled. Cancelling a cancelled job is a no-op and returns 200; a job that already finished returns 409.

Path
job_idREQUIRED
string
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/job/{job_id}/cancel \
  --request POST \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "status": "cancelled",
  "tasks_count": 50,
  "tasks_done": 32,
  "tasks_remaining": 0,
  "total_cost": 64,
  "refunded_tasks": 18,
  "refunded_credits": 18
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "JOB_NOT_FOUND",
  "message": "Job not found"
}
{
  "code": "JOB_NOT_CANCELLABLE",
  "message": "The job already finished and cannot be cancelled"
}