Search First, Crawl Second: A Source Discovery Workflow
Search first when the source URL is unknown or the task needs independent perspectives. Search should return candidates, not evidence. Normalize and shortlist those candidates, preserve domain and source-type diversity, then fetch only the pages needed to answer a defined question. Render when required content is missing from the initial HTML, and capture a screenshot only when visible state matters. If the URLs are already known and trusted, skip discovery and read them directly. The practical stop test is whether the retrieved sources close the evidence gaps you defined before searching.
First separate discovery from reading
Search and crawl solve different unknowns. Search asks, “Which pages might contain the answer?” Reading asks, “What does this selected page actually say, and can I verify the part I need?” A result title or snippet can support triage, but it is not a substitute for the source page.
Google's official description of Search stages separates URL discovery, crawling, indexing, and serving results. An agent workflow is smaller and task-specific, but the boundary is useful: finding a URL is not the same event as retrieving and validating its content.
In this article, “crawl second” means reading selected public pages with a one-page Fetch or Render request. It does not mean recursively traversing an entire domain.
| Tool stage | Use it when | Treat the output as | Failure signal | Next action |
|---|---|---|---|---|
| Search | The source URL is unknown, or the task needs alternative sources | Candidate URLs, titles, snippets, and ranking context | Results repeat one domain, miss the required source type, or only paraphrase one another | Reformulate the query or change the search channel |
| Fetch | A selected URL should expose the required content in HTML | Markdown, metadata, links, status, and URL identity | HTTP success but the required section, field, or phrase is absent | Check redirects and the extraction contract; consider Render only if JavaScript is the cause |
| Render | The required page state appears only after browser execution | Rendered Markdown and page fields | Timeout, blocked state, or still-missing required content | Record the failure; do not keep escalating without a new hypothesis |
| Screenshot | Layout, disclosure, visual comparison, or UI state is part of the evidence | A visible-state artifact tied to a URL and time | Broken capture, wrong viewport, consent overlay, or hidden target state | Correct the capture conditions or mark visual evidence unavailable |
This division keeps discovery broad enough to surface competing sources while keeping page access narrow, deliberate, and auditable.
Build a source-shortlist funnel
Start with an evidence contract, not a query. Write down the claim or decision you need to support, the preferred source type, the freshness requirement, and what would count as a contradiction. A query without that contract tends to produce a pile of topically related pages rather than a defensible answer.
The funnel has six states:
- Discover. Choose the search channel that matches the source: general pages for documentation and public web sources, news for recent coverage, scholar for papers, images for visual candidates, and video for talks or demonstrations.
- Normalize. Remove fragments and known tracking parameters, preserve the original URL, and resolve obvious URL variants without assuming that similar-looking paths are identical.
- Shortlist. Prefer pages that directly address the evidence need. Keep meaningful domain and source-type diversity instead of filling the list with syndicated copies or multiple pages from one publisher.
- Check access. Confirm that the page is public and that automated access is appropriate. RFC 9309 defines the Robots Exclusion Protocol for crawler access rules, while also making clear that robots rules are not access authorization.
- Read and validate. Fetch the smallest set of pages that can close the evidence gap. Require specific fields, headings, phrases, or metadata rather than accepting any non-empty response as success.
- Resolve or reopen. If the sources agree and cover the contract, stop. If they conflict, are stale, or leave a material gap, issue a narrower query that names the missing entity, date, version, or counterclaim.
The Web is connected by links, but a link only identifies a relationship and target; RFC 8288's Web Linking model does not make the linked resource authoritative. Treat links as routes to inspect, not endorsements.
Select for evidence value, not position alone
Use qualitative gates before inventing a score. A numeric ranking can hide weak assumptions when the team has not calibrated its weights.
| Selection signal | Question to ask | Keep the candidate when |
|---|---|---|
| Directness | Is this page the closest available source for the claim? | It is a standard, official document, original dataset, first-party specification, or direct statement relevant to the task |
| Freshness | Could the claim have changed? | The page gives a current version, date, or maintained status appropriate to the decision |
| Independence | Does this add a distinct evidence path? | It is not merely repeating or syndicating another shortlisted source |
| Access and completeness | Can the agent retrieve the required section and provenance? | The response exposes the needed content, status, and URL identity without bypassing controls |
| Diversity | Would a different domain or source type reveal a blind spot? | It adds a relevant perspective rather than variety for its own sake |
Canonical metadata can help identify duplicate URL variants, but it remains an input to verification. Google's canonicalization guidance describes canonical selection as choosing a representative URL and notes that a declared preference is a hint, not a rule. Preserve requested, final, and canonical URLs separately when provenance matters.
A tested shortlist-to-read example
The following TypeScript takes a candidate list produced by an earlier search step, removes duplicate hosts, and reads the first selected public page through AnyCrawler's no-key Fetch path. It deliberately does not pretend that the free endpoint performs Search. In an authenticated workflow, populate candidates from the Page Search API or another appropriate channel, then keep the same shortlist and validation gates.
type Candidate = { url: string; title: string };
const candidates: Candidate[] = [
{
url: "https://www.rfc-editor.org/info/rfc9309/",
title: "RFC 9309: Robots Exclusion Protocol",
},
{
url: "https://developers.google.com/search/docs/fundamentals/how-search-works",
title: "In-depth guide to how Google Search works",
},
{
url: "https://developers.google.com/search/docs/crawling-indexing/canonicalization?hl=en",
title: "What is URL canonicalization",
},
];
function normalizeUrl(value: string): string {
const url = new URL(value);
url.hash = "";
for (const key of [...url.searchParams.keys()]) {
if (key.startsWith("utm_")) url.searchParams.delete(key);
}
return url.toString();
}
const seenHosts = new Set<string>();
const shortlist = candidates
.map((candidate) => ({ ...candidate, url: normalizeUrl(candidate.url) }))
.filter((candidate) => {
const host = new URL(candidate.url).hostname;
if (seenHosts.has(host)) return false;
seenHosts.add(host);
return true;
});
const target = shortlist[0];
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", target.url);
const response = await fetch(endpoint, { headers: { accept: "application/json" } });
const payload = await response.json() as {
requested_url?: string;
final_url?: string | null;
status_code?: number;
credits_used?: number;
results?: { title?: string; markdown?: string };
};
const markdown = payload.results?.markdown ?? "";
const valid = response.ok
&& payload.status_code === 200
&& markdown.includes("Robots Exclusion Protocol");
console.log({
httpStatus: response.status,
requestedUrl: payload.requested_url,
finalUrl: payload.final_url ?? null,
pageStatus: payload.status_code,
title: payload.results?.title,
hasRequiredPhrase: valid,
creditsUsed: payload.credits_used ?? null,
});
if (!valid) throw new Error("Selected source did not satisfy the read contract");
This example was run on September 1, 2026, against the public RFC Editor page. The HTTP response and page status were both 200, the title matched RFC 9309, and the Markdown contained the required phrase. The response exposed the requested URL, status, timestamp, elapsed-time, title, and Markdown fields. It did not return credits_used, so the recorded value is “not returned,” not zero. One successful page read proves this contract for that input only; it does not establish universal coverage or performance.
For production discovery, start at the Search API hub and choose the source channel explicitly. Once a candidate is selected, use Fetch for HTML-first content, Render when the required content depends on JavaScript, or Screenshot when visible state belongs in the evidence record.
Route failures back to the right state
A reliable workflow does not treat every failure as a reason to crawl more pages. It returns to the state that introduced the uncertainty.
- A snippet appears to answer the question. Read the source page anyway. The snippet may be truncated, stale, or detached from its qualifying context.
- Several results repeat the same claim. Check whether they cite or syndicate one origin. Domain diversity is useful only when it produces independent evidence.
- The page is blocked or requires authentication. Record the access boundary and select a lawful public alternative. Do not treat robots rules, a search snippet, or a cached copy as permission to bypass access controls.
- Fetch returns a shell without the required content. Compare the response with the evidence contract. Escalate to Render only when browser execution is a plausible missing step, not simply because JavaScript exists on the page.
- The page loads but the claim is visual. Add a screenshot tied to the same URL and capture time; do not infer layout, disclosure prominence, or visible state from Markdown alone.
- Sources disagree. Open a gap search that includes the disputed entity, date, version, or measurement. Keep both claims in the ledger until a more direct source resolves them.
This state machine fits naturally inside a larger source-verifying research agent: discover → shortlist → access-check → read → validate → done, with gap-search returning to discovery only when the evidence contract remains open.
Where should an agent stop searching?
The difficult boundary is not how many results to collect; it is how much unresolved risk the task can tolerate. A documentation lookup may stop after one current primary source answers a narrow field question. A time-sensitive comparison may need independent confirmation and a visible-state artifact. A controversial claim may remain unresolved even after many reads.
Define the stop condition before the first query: required claims covered, material contradictions resolved or disclosed, source identity preserved, freshness acceptable, and access failures recorded. Then make the agent expose which condition remains false. This keeps “search more” from becoming an unbounded fallback while leaving room to reopen discovery when new evidence changes the decision.
Frequently asked questions
Should an agent always search before crawling a page?
No. Search first when the URL is unknown, when a known source may be stale, or when the task needs independent perspectives. If a user supplies a trusted, current URL and the task is to extract a specific section, go directly to Fetch or Render. Adding Search in that case can introduce irrelevant candidates and extra work. Keep the stages separate so the agent can skip discovery without weakening validation of the page it reads.
Why is a search result snippet not enough evidence?
A snippet is selected for result presentation, not for your evidence contract. It may omit qualifications, combine separated text, reflect an older indexed version, or point to a page that is now inaccessible. Use the snippet to decide whether a result deserves attention, then retrieve the page and validate the exact section or field you need. If the page cannot be read, record that limitation instead of promoting the snippet to source evidence.
How many sources should make the shortlist?
There is no universal count. The shortlist should be the smallest set that can cover the required claims, expose meaningful disagreement, and satisfy the task's freshness and independence needs. A narrow specification question may need one primary document. A comparison may need several distinct source types. Stop adding candidates when they no longer close a named evidence gap; do not use a fixed number as a substitute for coverage criteria.
When should a selected result use Fetch instead of Render?
Start with Fetch when the required content is present in the initial HTML and the response satisfies your extraction checks. Move to Render only when a specific required element depends on browser execution, client-side routing, or delayed page state. An HTTP 200 response alone is not success, but the presence of JavaScript alone is not a reason to render. Validate the target heading, field, phrase, or metadata before choosing the heavier path.
Can canonical URLs remove every duplicate source automatically?
No. Canonical metadata is useful evidence about a publisher's preferred representative URL, but it can be missing, incorrect, or only partially aligned with the content you retrieved. Preserve the requested URL, redirect destination, and canonical value separately. Combine those signals with content comparison and publisher identity. Do not merge two sources merely because their titles look similar, and do not assume different URLs are independent without checking their origin.
How can an agent prove that the workflow is complete?
Make completion a set of observable gates: every required claim has a readable source, the source directly supports the recorded statement, freshness is acceptable, material contradictions are resolved or disclosed, and failed access attempts remain visible. The agent should return both the evidence ledger and the still-open conditions. “No more results found” is not proof of completeness; it is only a search outcome that may require a better query or an explicit uncertainty.






