Scraping API
https://scraping-api.datafuel.ai/api/v1
Help centerOpenAPI spec

Tasks

Synchronous single-target scraping

GET/task

List your tasks

Your tasks newest first, including the tasks inside jobs and crawls. Each item carries the target url (unlocker and map tasks) but never the result: fetch that with GET /task/{task_id}. Pass next_cursor back as cursor for the next page; it is absent on the last page. Unknown status or type values match nothing.

Query
status
string
pendingprocessingcompletedfailed
type
string
unlockerllm_scrapingserpmapcrawl
job_id
string
Only the tasks of this job or crawl. Not a UUID answers 400 INVALID_QUERY_PARAM.
start_date
string
YYYY-MM-DD (UTC), created on or after this day. Wrong format answers 400 INVALID_DATE_FORMAT.
end_date
string
YYYY-MM-DD (UTC), created on or before this day (inclusive). Before start_date answers 400 INVALID_DATE_RANGE.
limit
integer
Items per page; values above 200 are capped. Negative or not a number answers 400 INVALID_QUERY_PARAM. Default: 50.
cursor
string
next_cursor of the previous page. A cursor that does not decode answers 400 INVALID_CURSOR.
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "tasks": [
    {
      "id": "7b0e8c1a-3f4d-4e59-9a61-2c8d5e7f9b10",
      "job_id": "1d2c3b4a-5e6f-4a7b-8c9d-0e1f2a3b4c5d",
      "type": "unlocker",
      "status": "completed",
      "url": "https://example.com/product/1",
      "credit_cost": 1,
      "created_at": "2026-09-29T10:00:00Z",
      "processed_at": "2026-09-29T10:00:02Z"
    },
    {
      "id": "9c1f2e3d-4b5a-4c6d-8e7f-a0b1c2d3e4f5",
      "job_id": null,
      "type": "llm_scraping",
      "status": "failed",
      "credit_cost": 100,
      "created_at": "2026-09-29T09:58:11Z",
      "failed_at": "2026-09-29T09:58:40Z"
    }
  ],
  "next_cursor": "MTc1OTEzOTQ5MTAwMDAwMDAwMHw5YzFmMmUzZC00YjVhLTRjNmQtOGU3Zi1hMGIxYzJkM2U0ZjU"
}
{
  "code": "INVALID_QUERY_PARAM",
  "message": "limit, page and job_id must be valid values"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
POST/tasktype: unlocker · synchronous

Unlocker

Executes one scrape and waits for completion. The connection stays open until the result is ready, so set your HTTP client timeout generously (120s+ recommended for js_rendering or llm_scraping workloads).

Every response carries a metadata envelope (status_code, final_url, redirected, credits_used, blocked, ...) next to result; see ResultMeta. Bodies above 50 MB are cut off (request engine) or fail (browser engine). A blocked target (403/429/503, anti-bot wall) fails the task and refunds it; blocked, protection and status_code tell you why.

Accepted type values: unlocker, llm_scraping, serp, map. crawl is rejected here with 400 UNSUPPORTED_TASK_TYPE; use POST /crawl.

If the task outlives the request (10 minutes) the answer is 202 TASK_STILL_PROCESSING with the task id: poll GET /task/{id}, or resend with the same Idempotency-Key.

Headers
Idempotency-Key
string
Client-chosen key, unique per task on your account. If the connection drops before you receive the response, resend the same request with the same key: the API attaches to the task already running (or returns its stored result) instead of creating and charging a second one. Reusing a key with a different request (type, proxy fields or attributes) is rejected with 422.
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
ISO 3166-1 alpha-2 country code. Used when attributes.proxy_country is not set. Ignored for serp.
proxy_city
string
proxy_state
string
proxy_asn
string
proxy_session_idDEPRECATED
string
Accepted but ignored at this level. Set attributes.proxy_session_id instead.
proxy_ttlDEPRECATED
integer
Accepted but ignored at this level. Set attributes.proxy_ttl instead.
Attributes
urlREQUIRED
string
method
string
Case-insensitive; anything else answers 400 INVALID_ATTRIBUTES. Default: "GET".
GETPOSTPUTDELETEPATCHHEADOPTIONS
body
string
Raw request body sent to the target with method POST, PUT or PATCH, on both engines (js_rendering on or off).
content_type
string
Content-Type header sent with body. A Content-Type in headers takes precedence.
headers
object
header_order
array<string>
cookie_string
string
user_agent_type
string
chromefirefoxsafariedge
user_agent
string
Explicit User-Agent, overrides user_agent_type
js_rendering
boolean
Render the page in a real browser (browser pricing, slower). Off by default: a plain HTTP fetch keeps the exact URL you asked for; sites that redirect filtered URLs to a canonical page or render listings client-side need it on. Check redirected / final_url in the response either way. Default: false.
js_instructions
array<object> | object
Browser actions run after load (js_rendering only). Recommended form: an array of single-action objects, run in order, where an action may repeat, e.g. [{"fill": ["input[name=q]", "laptops"]}, {"click": "button[type=submit]"}, {"wait_ms": 1000}]. The older object form keyed by action ({"fill": [...], "click": "..."}) is still accepted, but its order is not guaranteed and an action cannot repeat. The supported actions and their argument shapes are listed by GET /config/js-instructions; actions marked iframe also accept the iframe_ prefix. Anything else (a string, a number, an array element that is not a single-action object, an unknown action) answers 400 INVALID_ATTRIBUTES.
block_resource
string | array<string>
Resource types to block in the browser: one value as a string, or several as an array. An unknown value answers 400 INVALID_ATTRIBUTES.
wait_for_selector
string
js_rendering only: CSS selector to wait for before capturing
wait_for_selector_timeout_ms
integer
js_rendering only: timeout of wait_for_selector Default: 3000.
extract_selector
string
Extract only matching elements instead of formatting the whole page. Always a string: either a bare CSS selector (result key selector) or a JSON-encoded object of name → selector for several fields. A JSON object (not a string) answers 400 INVALID_ATTRIBUTES. Append @attr to read an attribute instead of the text. The result is {"data": {name: [values...]}} and result_format is forced to json.
extract_regex
string
Regex extraction on the raw HTML. Always a string: either a bare RE2 pattern (result key regex) or a JSON-encoded object of name → pattern. The first capture group is returned when present, otherwise the whole match; all matches are collected as a list.
result_format
string
html: raw page. markdown: cleaned text with a title/meta header, best for LLMs. json: schema.org / JSON-LD and embedded JSON objects. png / jpeg: full-page screenshot, base64; needs js_rendering: true (the request engine returns html instead). Default: "html".
htmlmarkdownjsonpngjpeg
result_template
string
Comma-separated built-in extractors returning JSON: images, links, headings, phone_numbers, emails, meta_tags, tables, schema_org, all
include_images
boolean
Markdown only. false drops ![](src) images (and the links that only wrapped them); signed image URLs often dominate the token count of a listing page. The MCP server defaults this to false. Default: true.
main_content_only
boolean
Markdown only. Always render just the <main> / <article> container when the page has one. Default: false.
result_use_ai
boolean
Post-process the raw result with an LLM
result_ai_format
object
JSON shape the AI output must follow
result_ai_prompt
string
ai_provider
string
openaianthropicgoogle
ai_model
string
ai_api_key
string
Bring-your-own-key for AI post-processing; stored encrypted
proxy_session_id
string
Sticky session. Requests with the same id reuse the same proxy exit.
proxy_ttl
integer
Sticky session lifetime in seconds
proxy_type
string
Filled from the top-level proxy_type. Set the plan at the top level; that is the one that is billed.
proxy_country
string
Overrides the top-level proxy_country
proxy_city
string
Overrides the top-level proxy_city
proxy_state
string
Overrides the top-level proxy_state
proxy_asn
string
Overrides the top-level proxy_asn
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "unlocker",
    "proxy_type": "Premium",
    "proxy_country": "US",
    "attributes": {
      "url": "https://example.com/product/123",
      "method": "GET",
      "js_rendering": true,
      "wait_for_selector": "#price",
      "result_format": "json"
    }
  }'
Response
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "status": "completed",
  "status_code": 200,
  "final_url": "https://example.com/product/123",
  "redirected": false,
  "credits_used": 10,
  "duration_ms": 5432,
  "blocked": false,
  "result": {
    "data": "# Product 123\n\nPrice: $19.99\n\nIn stock, ships in 2 days."
  }
}
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "code": "TASK_STILL_PROCESSING",
  "message": "The task is still being processed. Please try again."
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "FORBIDDEN",
  "message": "You do not have permission to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "CONCURRENCY_LIMIT_REACHED",
  "message": "You have reached your concurrency limit"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
POST/tasktype: llm_scraping · synchronous

LLM scraping

Executes one scrape and waits for completion. The connection stays open until the result is ready, so set your HTTP client timeout generously (120s+ recommended for js_rendering or llm_scraping workloads).

Every response carries a metadata envelope (status_code, final_url, redirected, credits_used, blocked, ...) next to result; see ResultMeta. Bodies above 50 MB are cut off (request engine) or fail (browser engine). A blocked target (403/429/503, anti-bot wall) fails the task and refunds it; blocked, protection and status_code tell you why.

Accepted type values: unlocker, llm_scraping, serp, map. crawl is rejected here with 400 UNSUPPORTED_TASK_TYPE; use POST /crawl.

If the task outlives the request (10 minutes) the answer is 202 TASK_STILL_PROCESSING with the task id: poll GET /task/{id}, or resend with the same Idempotency-Key.

Headers
Idempotency-Key
string
Client-chosen key, unique per task on your account. If the connection drops before you receive the response, resend the same request with the same key: the API attaches to the task already running (or returns its stored result) instead of creating and charging a second one. Reusing a key with a different request (type, proxy fields or attributes) is rejected with 422.
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
ISO 3166-1 alpha-2 country code. Used when attributes.proxy_country is not set. Ignored for serp.
proxy_city
string
proxy_state
string
proxy_asn
string
proxy_session_idDEPRECATED
string
Accepted but ignored at this level. Set attributes.proxy_session_id instead.
proxy_ttlDEPRECATED
integer
Accepted but ignored at this level. Set attributes.proxy_ttl instead.
Attributes
promptREQUIRED
string
engineREQUIRED
string
Every engine the API implements. Which ones accept work right now is in GET /config/capabilities.
openaigeminigoogle_ai_modeperplexitycopilot
websearch
boolean
Let the engine browse the web while answering
follow_up_prompt
string
Second prompt sent in the same conversation
proxy_country
string
Exit country; falls back to the top-level proxy_country
location
string
Location the engine should answer for (used by gemini and perplexity); defaults to the exit country
result_format
string
Ignored; LLM scraping always returns json.
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "llm_scraping",
    "attributes": {
      "prompt": "best CRM tools for startups 2026",
      "engine": "perplexity",
      "websearch": true
    }
  }'
Response
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "status": "completed",
  "credits_used": 100,
  "duration_ms": 18210,
  "result": {
    "data": {
      "answer": "For most startups HubSpot's free CRM is the easiest start; Pipedrive suits sales-led teams…",
      "sources": [
        {
          "title": "The 10 best CRM tools for startups",
          "url": "https://example.com/best-crm"
        }
      ]
    }
  }
}
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "code": "TASK_STILL_PROCESSING",
  "message": "The task is still being processed. Please try again."
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "FORBIDDEN",
  "message": "You do not have permission to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "CONCURRENCY_LIMIT_REACHED",
  "message": "You have reached your concurrency limit"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
POST/tasktype: serp · synchronous

SERP

Executes one scrape and waits for completion. The connection stays open until the result is ready, so set your HTTP client timeout generously (120s+ recommended for js_rendering or llm_scraping workloads).

Every response carries a metadata envelope (status_code, final_url, redirected, credits_used, blocked, ...) next to result; see ResultMeta. Bodies above 50 MB are cut off (request engine) or fail (browser engine). A blocked target (403/429/503, anti-bot wall) fails the task and refunds it; blocked, protection and status_code tell you why.

Accepted type values: unlocker, llm_scraping, serp, map. crawl is rejected here with 400 UNSUPPORTED_TASK_TYPE; use POST /crawl.

If the task outlives the request (10 minutes) the answer is 202 TASK_STILL_PROCESSING with the task id: poll GET /task/{id}, or resend with the same Idempotency-Key.

Headers
Idempotency-Key
string
Client-chosen key, unique per task on your account. If the connection drops before you receive the response, resend the same request with the same key: the API attaches to the task already running (or returns its stored result) instead of creating and charging a second one. Reusing a key with a different request (type, proxy fields or attributes) is rejected with 422.
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
ISO 3166-1 alpha-2 country code. Used when attributes.proxy_country is not set. Ignored for serp.
proxy_city
string
proxy_state
string
proxy_asn
string
proxy_session_idDEPRECATED
string
Accepted but ignored at this level. Set attributes.proxy_session_id instead.
proxy_ttlDEPRECATED
integer
Accepted but ignored at this level. Set attributes.proxy_ttl instead.
Attributes
queryREQUIRED
string
The search query (q). Accepts Google operators (site:, intitle:, ...).
country
string
Two-letter country for Google results (gl), e.g. us, gb, fr.
language
string
Language code (hl), e.g. en, en-gb, es-419.
page
integer
1-based result page; offset is (page-1) * 10 (start). Default: 1.
google_domain
string
Google domain to query, e.g. google.co.uk (a leading www. is dropped). Default google.com; pooled per domain.
location
string
Free-text origin of the search, sent to Google as uule. Exclusive with uule and lat/lon. Google may still place the local pack by the exit IP.
uule
string
Google-encoded location. Exclusive with location, lat/lon.
lat
number
GPS latitude of the search origin, -90 to 90. Requires lon.
lon
number
GPS longitude, -180 to 180. Requires lat.
radius
integer
Bias radius in meters, carried inside the GPS location. Needs one of location, uule or lat/lon.
cr
string
Limit to countries, |-delimited upper-case codes, e.g. countryFR|countryDE.
lr
string
Limit to languages, |-delimited, e.g. lang_fr|lang_de.
tbs
string
Advanced "to be searched" filters (dates, news, patents, ...).
safe
string
SafeSearch. active filters explicit content, off disables.
activeoff
nfpr
boolean
true excludes auto-corrected-query results (served via Google's own link).
filter
boolean
Similar/omitted-results filter. true on (default); sends 0 only when explicitly false.
uds
string
Google filter string, echoed back per filter in the results.
kgmid
string
Knowledge Graph listing id.
si
string
Cached/encrypted search-params blob (Knowledge Graph tabs).
ludocid
string
Google CID of a place.
lsig
string
Forces the knowledge-graph map view to render.
ibp
string
Layout/expansion renderer, e.g. gwp;0,7.
proxy_country
string
Accepted but not used yet. SERP searches exit through DataFuel's own pool; use country and language to target results.
result_format
string
json structured results, markdown LLM-optimized, html raw served page. Default: "json".
jsonmarkdownhtml
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "serp",
    "proxy_country": "US",
    "attributes": {
      "query": "best CRM tools for startups 2026",
      "country": "us",
      "language": "en",
      "page": 1,
      "result_format": "json"
    }
  }'
Response
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "status": "completed",
  "credits_used": 50,
  "result": {
    "data": {
      "query": "best crm for startups",
      "url": "https://www.google.com/search?q=best+crm+for+startups&hl=en&gl=us",
      "page": 1,
      "organic": [
        {
          "position": 1,
          "title": "The 10 best CRM tools for startups",
          "url": "https://example.com/best-crm",
          "displayed_link": "example.com › best-crm",
          "snippet": "We tested 25 CRMs…"
        }
      ],
      "questions": [
        {
          "text": "Which CRM is best for a startup?"
        }
      ],
      "related_searches": [
        {
          "query": "free crm for startups"
        }
      ]
    }
  }
}
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "code": "TASK_STILL_PROCESSING",
  "message": "The task is still being processed. Please try again."
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "FORBIDDEN",
  "message": "You do not have permission to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "CONCURRENCY_LIMIT_REACHED",
  "message": "You have reached your concurrency limit"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
POST/tasktype: map · synchronous

Map

Executes one scrape and waits for completion. The connection stays open until the result is ready, so set your HTTP client timeout generously (120s+ recommended for js_rendering or llm_scraping workloads).

Every response carries a metadata envelope (status_code, final_url, redirected, credits_used, blocked, ...) next to result; see ResultMeta. Bodies above 50 MB are cut off (request engine) or fail (browser engine). A blocked target (403/429/503, anti-bot wall) fails the task and refunds it; blocked, protection and status_code tell you why.

Accepted type values: unlocker, llm_scraping, serp, map. crawl is rejected here with 400 UNSUPPORTED_TASK_TYPE; use POST /crawl.

If the task outlives the request (10 minutes) the answer is 202 TASK_STILL_PROCESSING with the task id: poll GET /task/{id}, or resend with the same Idempotency-Key.

Headers
Idempotency-Key
string
Client-chosen key, unique per task on your account. If the connection drops before you receive the response, resend the same request with the same key: the API attaches to the task already running (or returns its stored result) instead of creating and charging a second one. Reusing a key with a different request (type, proxy fields or attributes) is rejected with 422.
Request fields
typeREQUIRED
string
Every registered task type. POST /task accepts unlocker, llm_scraping, serp and map (crawl answers 400 UNSUPPORTED_TASK_TYPE, use POST /crawl); POST /job accepts unlocker, llm_scraping and serp.
unlockerllm_scrapingserpmapcrawl
proxy_type
string
Ignored for llm_scraping and serp.
BasicPremium
proxy_country
string
ISO 3166-1 alpha-2 country code. Used when attributes.proxy_country is not set. Ignored for serp.
proxy_city
string
proxy_state
string
proxy_asn
string
proxy_session_idDEPRECATED
string
Accepted but ignored at this level. Set attributes.proxy_session_id instead.
proxy_ttlDEPRECATED
integer
Accepted but ignored at this level. Set attributes.proxy_ttl instead.
Attributes
urlREQUIRED
string
Page or site to map. Only http(s) with a host.
search
string
Keep only links whose URL or title contains this text (case-insensitive)
sitemap
string
Map only this sitemap (and its children) instead of discovering sitemaps and fetching the page. Must be on the same site as url. Take it from the sitemaps list of a previous result.
limit
integer
Maximum number of links to return. 0 or omitted means 5000; values above 10000 are capped.
include_subdomains
boolean
Also keep links on subdomains of the target host Default: false.
ignore_sitemap
boolean
Skip sitemap discovery; use only links found on the page Default: false.
sitemap_only
boolean
Skip the page fetch; use only sitemap entries. Cannot be combined with ignore_sitemap. Default: false.
user_agent_type
string
chromefirefoxsafariedge
user_agent
string
proxy_session_id
string
Sticky session. Requests with the same id reuse the same proxy exit.
proxy_ttl
integer
Sticky session lifetime in seconds
proxy_country
string
Overrides the top-level proxy_country
proxy_city
string
proxy_state
string
proxy_asn
string
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task \
  --request POST \
  --header 'Content-Type: application/json' \
  --header 'X-API-Key: df_key_your_key_here' \
  --data '{
    "type": "map",
    "attributes": {
      "url": "https://example.com",
      "search": "blog",
      "limit": 200
    }
  }'
Response
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "status": "completed",
  "credits_used": 1,
  "result": {
    "data": {
      "url": "https://example.com/",
      "links": [
        {
          "url": "https://example.com/blog/",
          "source": "sitemap",
          "lastmod": "2026-08-01"
        }
      ],
      "total": 1,
      "truncated": false,
      "sitemaps": [
        "https://example.com/sitemap.xml"
      ],
      "credits": 1,
      "page_status_code": 200
    }
  }
}
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "code": "TASK_STILL_PROCESSING",
  "message": "The task is still being processed. Please try again."
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "INSUFFICIENT_CREDITS",
  "message": "You do not have enough credits to perform this action"
}
{
  "code": "FORBIDDEN",
  "message": "You do not have permission to perform this action"
}
{
  "code": "TASK_ALREADY_EXISTS",
  "message": "A task with this id already exists; retry the request"
}
{
  "code": "IDEMPOTENCY_KEY_REUSED",
  "message": "Idempotency-Key was already used for a different request"
}
{
  "code": "CONCURRENCY_LIMIT_REACHED",
  "message": "You have reached your concurrency limit"
}
{
  "code": "ENGINE_UNAVAILABLE",
  "message": "This engine is switched off at the moment: temporarily unavailable"
}
GET/task/{task_id}

Get a task result

Same body as POST /task. While the task is pending or processing the answer is 202 TASK_STILL_PROCESSING (body {id, code, message}); poll again.

Path
task_idREQUIRED
string
Request
cURL
curl https://scraping-api.datafuel.ai/api/v1/task/{task_id} \
  --header 'X-API-Key: df_key_your_key_here'
Response
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "status": "completed",
  "status_code": 200,
  "final_url": "https://example.com/product/123",
  "redirected": false,
  "credits_used": 10,
  "duration_ms": 5432,
  "blocked": false,
  "result": {
    "data": "# Product 123\n\nPrice: $19.99\n\nIn stock, ships in 2 days."
  }
}
{
  "id": "123e4567-e89b-12d3-a456-426614174000",
  "code": "TASK_STILL_PROCESSING",
  "message": "The task is still being processed. Please try again."
}
{
  "code": "INVALID_ATTRIBUTES",
  "message": "Invalid attributes for selected task type"
}
{
  "code": "UNAUTHORIZED",
  "message": "You are not authorized to perform this action"
}
{
  "code": "JOB_NOT_FOUND",
  "message": "Job not found"
}