Scraping API
Help center

Scrape a list of URLs

Queue many known URLs in one job, poll its progress and collect the per-URL results.

When you already know which pages you want, POST /job queues all of them at once. It returns an id immediately; you poll for progress and fetch results when it finishes. Each URL costs the same as a single task.

Queue the job

Shell
curl https://scraping-api.datafuel.ai/api/v1/job \
  --request POST \
  --header "Content-Type: application/json" \
  --header "X-API-Key: df_key_your_key_here" \
  --header "Idempotency-Key: catalog-import-2026-09-18" \
  --data '{
    "type": "unlocker",
    "multithreaded": true,
    "proxy_type": "Basic",
    "attributes": {
      "urls": [
        "https://example.com/p/1",
        "https://example.com/p/2",
        "https://example.com/p/3"
      ],
      "result_format": "markdown",
      "main_content_only": true
    }
  }'
JSON
{ "id": "9f1d3c2a-4b6e-4f8a-9c7d-2e5b8a1f0d43" }

The attributes are the same as for a single task, with urls (plural) in place of url. For llm_scraping jobs use prompts, for serp jobs queries. A job needs at least two targets; a single one is rejected with 400 JOB_REQUIRES_MULTIPLE_TARGETS, so send it to POST /task instead. multithreaded: true runs the URLs in parallel up to your account’s concurrency limit; leave it off to fetch them one by one.

Poll until it finishes

Code
GET /job/{id}

The response carries status (pending, processing, completed, completed_with_errors, failed or cancelled) and the counters tasks_count, tasks_done and tasks_remaining. Poll every two seconds while the status is pending or processing. There are no webhooks yet.

Collect the results

Code
GET /job/{id}/results

Returns every task of the job in one response. Each entry has the same shape as a single task result: the page content under result.data and the metadata envelope next to it. A failed URL shows up with status: failed, credits_used: 0 and an error string; it does not fail the whole job.

Cancel a job

Code
POST /job/{id}/cancel

Tasks that have not started yet are refunded straight away; tasks already running finish and are charged as usual. The job ends with status cancelled and the response says how many tasks and credits were refunded. Cancelling a finished job returns 409 JOB_NOT_CANCELLABLE.

Billing

Each URL is charged when it is queued and refunded if it fails. Blocked pages count as failed. The job’s total_cost is the sum charged at queue time and does not subtract refunds. GET /users/@me/balance shows the net effect.

Job or crawl?

A job fetches exactly the URLs you give it. A crawl discovers URLs by following links. If you can list the URLs up front, a job is cheaper and its cost is known before you start. If you only have a start page, see Crawl a site, or run Map the URLs of a site first and feed its output into a job.