Map the URLs of a site
Get a deduplicated list of a site's URLs for one credit, before you decide what to scrape.
POST /map reads a site’s sitemaps and the links on the start page and returns one deduplicated list of same-site URLs. Nothing is scraped and the price is flat, so it is the cheapest way to see how big a section is before you spend credits on a crawl or a job.
Request
curl https://scraping-api.datafuel.ai/api/v1/map \
--request POST \
--header "Content-Type: application/json" \
--header "X-API-Key: df_key_your_key_here" \
--data '{
"attributes": {
"url": "https://example.com/docs",
"search": "guide",
"limit": 500
}
}'
The call is synchronous and returns the list in the response, under result.data:
{
"id": "…",
"credits_used": 1,
"result": {
"data": {
"url": "https://example.com/docs",
"links": [{ "url": "https://example.com/docs/guide", "source": "sitemap", "lastmod": "2026-08-01" }],
"total": 1,
"truncated": false,
"sitemaps": ["https://example.com/sitemap.xml"],
"credits": 1,
"page_status_code": 200
}
}
}
| Attribute | Effect |
|---|---|
url | Start page. Sitemaps are looked up from the site root. |
search | Keep only links whose URL or title contains this text (case-insensitive). |
limit | Maximum number of URLs returned. Default 5 000, capped at 10 000. |
sitemap | Sitemap URL to use instead of the discovered ones. Must be on the same site. |
sitemap_only | Skip the start page, use sitemaps only. Cannot be combined with ignore_sitemap. |
ignore_sitemap | Skip sitemaps, use links on the start page only. |
include_subdomains | Also return URLs on subdomains of the start host. |
proxy_type and proxy_country work as on any other request.
Price
One credit per call on Basic, ten on Premium, regardless of how many links come back.
When the list is empty
When links is empty, reason says why: page_blocked, page_error, page_unreachable, no_links_on_page, no_sitemap, sitemap_no_entries or search_no_match. no_links_on_page usually means the page renders its navigation client-side. Map always fetches the page without a browser, so scrape that page with POST /task and js_rendering: true to get its links, or point sitemap at a known sitemap URL.
What to do with the list
- Feed it into
POST /jobto scrape exactly those pages. See Scrape a list of URLs. - Use it to choose
include_pathsand a realisticmax_pagesfor a crawl. See Crawl a site.