Scraping API
Help center

Map the URLs of a site

Get a deduplicated list of a site's URLs for one credit, before you decide what to scrape.

POST /map reads a site’s sitemaps and the links on the start page and returns one deduplicated list of same-site URLs. Nothing is scraped and the price is flat, so it is the cheapest way to see how big a section is before you spend credits on a crawl or a job.

Request

Shell
curl https://scraping-api.datafuel.ai/api/v1/map \
  --request POST \
  --header "Content-Type: application/json" \
  --header "X-API-Key: df_key_your_key_here" \
  --data '{
    "attributes": {
      "url": "https://example.com/docs",
      "search": "guide",
      "limit": 500
    }
  }'

The call is synchronous and returns the list in the response, under result.data:

JSON
{
  "id": "…",
  "credits_used": 1,
  "result": {
    "data": {
      "url": "https://example.com/docs",
      "links": [{ "url": "https://example.com/docs/guide", "source": "sitemap", "lastmod": "2026-08-01" }],
      "total": 1,
      "truncated": false,
      "sitemaps": ["https://example.com/sitemap.xml"],
      "credits": 1,
      "page_status_code": 200
    }
  }
}
AttributeEffect
urlStart page. Sitemaps are looked up from the site root.
searchKeep only links whose URL or title contains this text (case-insensitive).
limitMaximum number of URLs returned. Default 5 000, capped at 10 000.
sitemapSitemap URL to use instead of the discovered ones. Must be on the same site.
sitemap_onlySkip the start page, use sitemaps only. Cannot be combined with ignore_sitemap.
ignore_sitemapSkip sitemaps, use links on the start page only.
include_subdomainsAlso return URLs on subdomains of the start host.

proxy_type and proxy_country work as on any other request.

Price

One credit per call on Basic, ten on Premium, regardless of how many links come back.

When the list is empty

When links is empty, reason says why: page_blocked, page_error, page_unreachable, no_links_on_page, no_sitemap, sitemap_no_entries or search_no_match. no_links_on_page usually means the page renders its navigation client-side. Map always fetches the page without a browser, so scrape that page with POST /task and js_rendering: true to get its links, or point sitemap at a known sitemap URL.

What to do with the list

  • Feed it into POST /job to scrape exactly those pages. See Scrape a list of URLs.
  • Use it to choose include_paths and a realistic max_pages for a crawl. See Crawl a site.