If you already have the page URL, use a page Fetch API to read that URL and validate the returned content. If you need candidate sources but do not know their URLs, use Search first and fetch only the useful results. A web crawler API becomes the right tool when the task requires discovering and visiting a bounded set of linked pages across a site. These are different scopes: returning links from one page does not mean the service followed them. The choice depends on the input you have, the coverage you need, and who will manage the queue of URLs.
The boundary is how many URLs the tool must discover
“Crawl” can describe a single page read or a traversal of a site, so the endpoint name alone is a poor specification. Ask what happens after the first response. Does the tool stop with one document, return candidate URLs for you to select, or schedule further requests under a defined scope?
| Job | Starting input | Useful output | State your application must own |
|---|---|---|---|
| Page Fetch | One known URL | Content and identity of that page | Whether the returned body satisfies the task |
| Page Search | A query with unknown source URLs | A shortlist of candidate pages | Selection, deduplication, and which candidates to read |
| Site crawl | A seed URL and site boundaries | Visited pages plus skipped and failed URLs | Scope rules, coverage, budgets, and job progress unless the crawler explicitly provides them |
A site crawl follows discovered URLs. For a concrete example of a service that offers that behavior, Cloudflare's crawl endpoint documentation describes following links from a starting URL with configurable depth and page limits. That is an example of a full crawl contract, not a claim about AnyCrawler's current page endpoint. AnyCrawler's published routes document Search, single-page Fetch and Render, and Screenshot; they do not establish an automatic recursive site-crawl feature.
Route the task you actually have
When a support ticket includes an exact documentation URL, start with AnyCrawler's Fetch page API. It accepts a URL with method: "fetch" and returns normalized page content, status, URL identity, and optional links. Check for the passage or field the ticket needs. If the useful content appears only after JavaScript runs, the Render page path uses the same page endpoint with browser execution. Neither path implies that linked pages were visited.
When an agent knows a topic but has no source URL, AnyCrawler's page Search API returns candidates for a query. Select sources, then read the relevant pages. Search results may suggest a URL; they do not prove the page's full content or map every page under a host. The application still decides which results deserve a page request.
For a task such as indexing every eligible page in one documentation section, define the allowed host and paths, decide how far links can be followed, and keep a visited set and a pending queue. A page API can be one worker in that design, but the traversal is your application logic unless a separate crawl service explicitly owns it. A link list on the first page is a set of leads, not evidence of completed coverage.
Test the one-page contract before building a queue
This credential-free probe uses AnyCrawler's public Free endpoint against https://example.com/. It is a small way to inspect a single-page response; the Free route is not a substitute for the authenticated POST /v1/crawl/page contract. The check requires the expected page text instead of treating a successful HTTP response as complete extraction.
const target = "https://example.com/";
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", target);
const response = await fetch(endpoint);
if (!response.ok) throw new Error(`Gateway HTTP ${response.status}`);
const data = await response.json();
if (!data.ok || data.status_code !== 200 ||
!data.results?.markdown?.includes("Example Domain")) {
throw new Error("Expected page content missing");
}
console.log({
requested_url: data.requested_url,
final_url: data.final_url ?? null,
status_code: data.status_code,
title: data.results.title,
credits_used: data.credits_used ?? null,
});
In the checked response, the gateway and target status were both 200, the title was Example Domain, and the expected text appeared in Markdown. The Free response did not include final_url or credits_used, so the example leaves those values null. It says nothing about recursive discovery or paid API usage. Record the actual fields your route returns rather than filling absent provenance or billing fields with guesses.
A real crawl needs a coverage contract
A crawler must decide which discovered URLs are eligible, normalize and deduplicate them, enforce limits, record redirects, and keep failures separate from successful reads. A queue that follows every link can drift into search pages, calendars, external sites, or repeated query variants. Define host and path boundaries before the first fetch. Respect the target site's published access rules; Google's robots.txt guide explains the role and limits of robots instructions for search crawlers. Site permission and appropriate request rates still need separate judgment.
Report visited, skipped, failed, and still-pending URLs. A “done” status should mean the eligible queue was exhausted within the chosen bounds, not that every page on the site was discovered. Pages absent from links or outside the crawl depth remain outside that result. If the goal is only to answer a question from a handful of known sources, this queue and coverage accounting may add work without improving the answer.
Verify the result against the original request
For a one-page read, compare the returned page identity and required passage with the requested URL. Keep redirects and missing fields visible. For Search, open the shortlisted pages before citing them. For a site crawl, inspect the visited set, exclusions, failures, and remaining queue against the declared host, path, and depth rules. Sample important pages for extraction completeness; a successful crawl job can still contain empty or irrelevant bodies.
When would the boundary change?
A project may start with a few exact URLs and later need reliable section-wide coverage. That change calls for a new question: should the application own traversal and use page reads as workers, or should a service own discovery, limits, and job reporting? Check the current endpoint contract and a small representative site before assuming either design. The deciding variable is the required coverage and evidence of completion, not the presence of “crawler” in a product name.
Frequently asked questions
Does AnyCrawler's page API crawl an entire website?
The published POST /v1/crawl/page contract describes a request for one URL, using Fetch or Render to return that page's content. Optional links can help an application plan follow-up requests, but the response does not establish that those links were visited. If the task needs site-wide discovery, define a separate traversal queue and its boundaries, or verify that a dedicated crawl service explicitly provides recursive visits and job-level coverage.
Does include_links turn a page Fetch into a crawler?
No. include_links asks the page request to return links found during that page read. Those links are candidates, and some may be irrelevant, external, duplicated, or outside the intended scope. A crawler would need to select eligible candidates, fetch them, find more links, track visited URLs, and stop under stated limits. Treating the first link list as a finished crawl would hide all of those decisions and failures.
If I do not know the URL, should I use Search or a site crawler?
Use Search when you have a topic or question and need candidate source pages, potentially across sites. Inspect the results and fetch the pages that answer the task. Use a bounded site crawler when the requirement is to traverse pages within a specified site or section. Search is a shortlist, not exhaustive coverage; a crawler is a traversal with scope and progress. The choice follows the source boundary in the assignment.
When should I render a page instead of fetching it?
Start with Fetch when the required text is present in the returned HTML-derived content. Check an expected heading, field, or passage rather than trusting status alone. Render is appropriate when JavaScript or browser behavior is needed before that content appears. The Render path changes how one page is read; it does not add automatic recursive discovery. If several pages need rendering, the application still needs a bounded list or queue of URLs.
What prevents a site crawl from expanding indefinitely?
Set eligibility rules for hosts and paths, a link depth, a page or request budget, and a rule for duplicate URLs. Keep visited, skipped, failed, and pending records so the job can explain why it stopped. Redirects and query variants deserve explicit handling because they can multiply apparent pages. A crawl should report the scope it covered and the scope it excluded; no finite job can prove it discovered every possible URL.
How do I know whether the chosen workflow succeeded?
Use the original task as the acceptance test. For a page read, check the requested and returned URL identity plus the exact content needed. For Search, read selected results before making claims from them. For a site crawl, compare visited and excluded URLs with the planned site boundary, review failures and remaining work, and sample extracted content. A 200 response or a finished job alone does not establish that the information is complete.






