Crawl
Asynchronous site crawl (dynamic job)
/crawl Crawl a website (asynchronous)
Starts a crawl: the start URL is fetched, the links found on it are followed
breadth-first, and every page is scraped with the unlocker options given in
attributes (result_format, extract_selector, headers, ...).
A crawl is a dynamic job of type: crawl: it returns immediately with a
job_id, grows its own task set while it runs, and stops when max_pages
pages have been queued, no link within max_depth is left, or credits run out
(stop_reason). Poll GET /crawl/{job_id} for progress and page through
GET /crawl/{job_id}/results. The same id also works with the generic
GET /job/{job_id} routes.
Only same-site links are followed (include_subdomains widens that to
subdomains). With allow_backward_links: false (default) only URLs under the
start URL's path are followed, so a crawl of /docs never wanders into /blog.
include_paths / exclude_paths are RE2 regular expressions matched against
path?query; exclude wins.
Pricing: each page is a regular unlocker task, charged when the page is
queued: 1 credit times the proxy plan's request multiplier, or the browser
multiplier when js_rendering is true. Failed pages, blocked ones included,
are refunded like any other task, and the links of a block page are never
followed (a blocked start URL is a free zero-page crawl). total_cost is the
sum charged at queue time and does not subtract refunds; sum credits_used
over the results for the net figure.
js_rendering: true renders every page in a real browser and follows the
links of the rendered DOM, so client-side rendered sites crawl correctly.
js_instructions, wait_for_selector and block_resource apply per page.
Not yet supported: result_use_ai, sitemap seeding (ignore_sitemap is
accepted and ignored), robots.txt, webhooks.
Idempotency-Keystringtypestringcrawl.proxy_typestring400 INVALID_PROXY_TYPE.proxy_countrystringproxy_citystringproxy_statestringproxy_asnstringattributesREQUIREDobjectresult_use_ai is rejected.curl https://scraping-api.datafuel.ai/api/v1/crawl \ --request POST \ --header 'Content-Type: application/json' \ --header 'X-API-Key: df_key_your_key_here' \ --data '{ "proxy_type": "Basic", "attributes": { "url": "https://example.com/docs", "max_pages": 100, "max_depth": 3, "exclude_paths": [ "\\.pdf$" ], "result_format": "markdown" } }'
{ "job_id": "d3f8ae31-aed3-4b39-a72f-efaa895533ff" }
{ "code": "INVALID_ATTRIBUTES", "message": "Invalid attributes for selected task type" }
{ "code": "UNAUTHORIZED", "message": "You are not authorized to perform this action" }
{ "code": "INSUFFICIENT_CREDITS", "message": "You do not have enough credits to perform this action" }
{ "code": "TASK_ALREADY_EXISTS", "message": "A task with this id already exists; retry the request" }
{ "code": "IDEMPOTENCY_KEY_REUSED", "message": "Idempotency-Key was already used for a different request" }
{ "code": "ENGINE_UNAVAILABLE", "message": "This engine is switched off at the moment: temporarily unavailable" }
/crawl/{job_id} Get crawl progress
job_idREQUIREDstringcurl https://scraping-api.datafuel.ai/api/v1/crawl/{job_id} \ --header 'X-API-Key: df_key_your_key_here'
{ "status": "completed", "stop_reason": "max_pages", "pages": { "discovered": 98, "enqueued": 30, "done": 30, "failed": 0, "skipped": 68 }, "depth_reached": 1, "total_cost": 30, "created_at": "2026-09-10T17:20:11Z", "updated_at": "2026-09-10T17:20:52Z" }
{ "code": "UNAUTHORIZED", "message": "You are not authorized to perform this action" }
{ "code": "JOB_NOT_FOUND", "message": "Job not found" }
/crawl/{job_id}/results Get crawl results (paginated)
Pages in discovery order, oldest first. Pages that are still pending or
processing are included as status stubs so a client can stream results while
the crawl runs. Pass next_cursor back as cursor to continue; it is absent
on the last page.
job_idREQUIREDstringcursorstringlimitinteger400 INVALID_REQUEST_BODY. Default: 100.curl https://scraping-api.datafuel.ai/api/v1/crawl/{job_id}/results \ --header 'X-API-Key: df_key_your_key_here'
{ "pages": [ { "url": "https://example.com/docs", "depth": 0, "task_id": "17777d82-5c1e-4b8a-9f3d-2a6b7c8d9e0f", "status": "completed", "status_code": 200, "credits_used": 1, "result": { "data": "- Meta: charset: UTF-8\n- Title: Docs\n\n# Docs" } }, { "url": "https://example.com/docs/intro", "depth": 1, "task_id": "05e8b11d-3a2f-4c6e-8d1b-7e9f0a1b2c3d", "status": "processing", "credits_used": 0, "result": { "status": "processing" } } ], "next_cursor": "MTc4OTA2MzQ5MTAwMDAwMDAwMHwwNWU4YjExZC0" }
{ "code": "INVALID_ATTRIBUTES", "message": "Invalid attributes for selected task type" }
{ "code": "UNAUTHORIZED", "message": "You are not authorized to perform this action" }
{ "code": "JOB_NOT_FOUND", "message": "Job not found" }
/crawl/{job_id}/cancel Cancel a crawl
Stops the crawl: no more links are followed. Tasks that have not started yet fail and their credits are
refunded right away; tasks already running finish and are billed as usual.
The crawl ends with status cancelled. Cancelling a cancelled crawl is a no-op
and returns 200; a crawl that already finished returns 409.
job_idREQUIREDstringcurl https://scraping-api.datafuel.ai/api/v1/crawl/{job_id}/cancel \ --request POST \ --header 'X-API-Key: df_key_your_key_here'
{ "status": "cancelled", "tasks_count": 40, "tasks_done": 25, "tasks_remaining": 0, "total_cost": 40, "refunded_tasks": 15, "refunded_credits": 15 }
{ "code": "UNAUTHORIZED", "message": "You are not authorized to perform this action" }
{ "code": "JOB_NOT_FOUND", "message": "Job not found" }
{ "code": "JOB_NOT_CANCELLABLE", "message": "The job already finished and cannot be cancelled" }