How to Build a Deep Research Agent That Verifies Sources
A deep research agent should not treat web search as the answer. It should turn the user’s request into explicit subquestions, discover candidate sources, read only the strongest candidates, record source and access state, map every material claim to evidence, search again for gaps or contradictions, and stop when the evidence contract is satisfied or the budget is exhausted. Search discovers; Fetch or Render reads; citation validation checks whether the cited passage actually supports the claim. The system is reliable only when it can expose both supported conclusions and unresolved gaps.
Define the research contract before the first query
An open-ended request such as “research this market” leaves the agent to invent scope, dates, jurisdictions, source standards, and an output format. That ambiguity becomes invisible once a polished answer appears. Convert the request into a research contract first:
- the decision or question the report must support;
- the target region, time window, entities, and exclusions;
- the claims that require primary or official sources;
- acceptable source types and freshness rules;
- the required output fields, citation style, and unresolved-gap format;
- limits for queries, sources, tool calls, elapsed time, and credits.
This is not paperwork around the agent. It determines whether the run has a measurable finish line. OpenAI’s current deep research guidance makes the same implementation pressure visible in a product-specific way: the research step expects a sufficiently formed prompt, exposes tool actions and inline citations, and allows a caller to constrain tool calls. Your stack may use different models and tools, but scope, observable actions, and budgets remain useful control points.
Use a state machine, not one long prompt
Treat deep research as a loop with explicit transitions:
- Frame — decompose the contract into answerable subquestions.
- Discover — search for candidate sources by subquestion and source type.
- Select — rank for authority, directness, freshness, independence, and access.
- Read — Fetch known HTML-first pages; Render only when required content depends on JavaScript.
- Verify — map each material claim to a passage, table row, metadata field, or other retrievable evidence.
- Gap search — create follow-up queries for missing, conflicting, or stale evidence.
- Stop — return the report when the evidence contract passes, or return a bounded partial result when the budget ends.
The important transition is from Verify back to Gap search. A linear search-and-summarize pipeline cannot distinguish “the answer is complete” from “the agent stopped looking.” A state machine can attach a reason to each new query and a reason to the final stop.
Choose the access path by the evidence you need
The search result is a candidate record. The cited source is the document you opened and checked. Keep those two layers separate.
| Need | Use | Evidence to retain | Do not assume |
|---|---|---|---|
| The relevant URL is unknown | Search | query, channel, rank, result URL, title, snippet | the snippet proves the final claim |
| A known page exposes useful HTML | Fetch | requested URL, final URL, status, Markdown, access time | HTTP 200 means the required passage exists |
| Required content appears after JavaScript | Render | wait policy, final URL, status, rendered Markdown, access time | every dynamic page needs a browser |
| A paper or DOI needs identity checks | Structured registry | DOI, title, authors, venue, dates, record source | metadata provides the full text or validates its conclusions |
| Layout or visible state is material | Screenshot | URL, capture time, viewport, image key or hash | an image captures every semantic detail |
For scholarly work, the Crossref REST API documentation is a useful example of an identity layer: it exposes member-deposited bibliographic metadata and DOI lookup in JSON. That can help an agent verify which work it found, but metadata alone does not show that a paper supports a scientific claim.
If you are implementing this flow with AnyCrawler, start at the Search API hub to discover candidates, then use Fetch or Render only for selected sources. The AI web scraping decision guide is the next decision point when you need to separate reading a page from interacting with one.
Make the source ledger the system of record
A source list is not enough. Store the relationship between a query, a retrieved document, and the claim it is supposed to support.
type EvidenceState = "supported" | "contradicted" | "gap" | "inaccessible";
type SourceLedgerEntry = {
query: string;
requestedUrl: string;
finalUrl: string | null;
canonicalUrl: string | null;
accessStatus: number | null;
accessedAt: string;
sourceType: "official" | "standard" | "paper" | "registry" | "report" | "other";
title: string | null;
contentHash: string | null;
claim: string;
citationLocation: { field: string; text: string } | null;
evidenceState: EvidenceState;
failureCode: string | null;
};
Do not collapse requested, final, and canonical URLs into one field. The requested URL explains what the agent tried. The final URL captures redirects. The canonical URL can help with identity and deduplication, but it is a page assertion that still needs review. A content hash helps identify an exact retrieved body; it does not prove that two different pages make the same claim.
The ledger should also preserve discovery sources that never become citations. OpenAI’s web search documentation distinguishes inline citation annotations from a broader sources field containing consulted URLs. That exact response shape is vendor-specific, but the distinction is general: “considered” and “used as evidence” are different sets.
Run and test one retrieval step before adding autonomy
The following Node.js example starts after source selection. It uses AnyCrawler’s public no-key endpoint to retrieve one permitted public URL, independently follows redirects to record the final URL, and marks one claim as supported only when the expected passage is present.
const query = "What is the purpose of the Example Domain?";
const sourceUrl = "https://example.com/";
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", sourceUrl);
const [pageResponse, crawlResponse] = await Promise.all([
fetch(sourceUrl, { method: "HEAD", redirect: "follow" }),
fetch(endpoint, { headers: { accept: "application/json" } }),
]);
if (!crawlResponse.ok) {
throw new Error(`AnyCrawler request failed: ${crawlResponse.status}`);
}
const data = await crawlResponse.json();
const markdown = data.results?.markdown ?? "";
const claim = "Example Domain is reserved for documentation examples.";
const evidence = "This domain is for use in documentation examples";
const ledger = {
query,
requestedUrl: data.requested_url,
finalUrl: pageResponse.url,
accessStatus: data.status_code,
accessedAt: data.timestamp,
title: data.results?.title ?? null,
claim,
citationLocation: markdown.includes(evidence)
? { field: "results.markdown", text: evidence }
: null,
evidenceState: markdown.includes(evidence) ? "supported" : "gap",
};
console.log(JSON.stringify(ledger, null, 2));
This exact example was run on August 31, 2026 UTC. The AnyCrawler endpoint returned HTTP 200, target status 200, title Example Domain, and Markdown containing the evidence phrase. The independent request ended at https://example.com/, so the ledger state was supported. The public response did not include a credits field, and no authenticated Search, Fetch, or Render request was claimed. The observed timing belongs only to that single request and is not a performance benchmark. You can reproduce the retrieval path in the free crawler before wiring an authenticated production workflow.
Validate citations at the claim level
A URL attached to a paragraph is not yet a verified citation. For each material claim, run four checks:
- Entailment: does the quoted or structured evidence support the sentence as written?
- Scope: does the claim preserve the source’s geography, dates, sample, version, and qualifiers?
- Identity: is this the intended document, edition, paper, organization, or dataset?
- Access state: did the agent read the source successfully, or only see a snippet, cached copy, metadata record, or error page?
Record contradicted when a source directly conflicts with the claim and gap when evidence is missing. Do not silently downgrade either state to “probably true.” If two credible sources disagree, preserve both claims with their conditions and open a focused gap query.
Access policy belongs in the same workflow. The Robots Exclusion Protocol defines how service owners publish crawler rules and explicitly notes that those rules are not access authorization. A research agent should honor applicable crawl rules and still treat authentication, terms, privacy, copyright, and organizational policy as separate boundaries. It should never interpret a failed fetch as permission to bypass the source.
Untrusted page text is data, not an instruction stream. Keep retrieved content outside system and tool-control instructions, validate tool arguments, restrict outbound destinations, and never allow a page to cause private context to be sent to a new URL. This matters most when a run mixes public web research with private connectors or internal documents.
Classify failures before retrying
Retries should repair transient transport conditions, not repeat a bad plan.
| Failure class | Examples | Next transition |
|---|---|---|
| Retryable transport | connection reset, temporary gateway failure, server asks the client to wait | bounded retry, then inaccessible |
| Permanent request | invalid URL, unsupported scheme, authentication required | correct input or stop |
| Content invalid | status is successful but required passage or field is absent | Render, alternate source, or gap search |
| Identity conflict | redirect or canonical points to the wrong work | reject candidate and search again |
| Evidence conflict | credible sources support incompatible conclusions | narrow the claim and preserve both conditions |
| Budget exhausted | query, source, credit, or time limit reached | return a partial report with gaps |
Use server signals where available. HTTP Semantics defines status meanings and the Retry-After field used by servers to suggest when a client should try again. Your adapter should add API-specific rules, but a fixed “retry everything three times” policy can waste budget and repeat permanent failures.
Stop on evidence coverage, not source count
A useful stop evaluator runs after every verification pass. Stop with complete only when all of these are true:
- every required subquestion is answered or explicitly marked out of scope;
- every material factual claim has at least one retrievable citation location;
- high-impact or disputed claims meet the contract’s corroboration rule;
- contradictions are resolved, scoped, or exposed to the reader;
- source freshness and identity checks pass;
- no open gap is likely to change the decision the report supports.
Stop with partial when the budget is exhausted, a required source remains inaccessible, or a contradiction cannot be resolved. The result should list missing evidence and the next query that would reduce uncertainty. “Partial” is a valid research outcome; hiding the gap behind fluent prose is not.
The final report should be generated from the ledger, not from an untraceable conversation buffer. Include claim-to-source mappings, accessed times, unresolved gaps, and the exact scope. Keep an audit record of queries and tool outcomes separately from the reader-facing answer so the interface stays useful without losing provenance.
The open boundary: confidence is not completeness
Even a well-instrumented deep research agent cannot prove that it found every relevant source. Search indexes change, private and paywalled material may remain inaccessible, pages are revised, and two sources can share the same upstream error. More citations can increase the appearance of certainty without adding independent evidence.
The next design question is therefore not “How many sources are enough?” It is “What evidence would change this decision, and did the agent have a fair chance to find it?” Teams should define that answer per use case. A product comparison, literature review, incident analysis, and regulatory memo need different freshness, independence, and human-review gates. The ledger makes those policies enforceable, but it cannot choose the acceptable risk on its own.
Frequently asked questions
What is a deep research agent?
A deep research agent is a system that decomposes a question, searches for candidate sources, reads selected documents, checks claims against evidence, and repeats targeted searches until a stopping rule is met. It differs from a one-shot search summary because it preserves source state and unresolved gaps. The useful unit is not the number of pages collected; it is a claim linked to retrievable evidence with clear scope, identity, and access status.
Should a research agent cite search result snippets?
Treat snippets as discovery context, not final evidence. They may be truncated, stale, assembled from a page section that changes, or missing the qualifications surrounding a sentence. The agent should open the selected result, record the final URL and status, retrieve the relevant content, and save a citation location. If the source cannot be read, label it inaccessible or a gap instead of citing the snippet as though the document was verified.
When should a deep research agent use Fetch instead of Render?
Use Fetch when the required article, documentation, or reference content is already present in the returned HTML. It is the simpler reading path and should be validated by checking for expected passages or fields. Upgrade to Render only when JavaScript creates the content or navigation state needed for the research claim. The presence of JavaScript alone is not sufficient; the decision depends on whether the evidence contract fails without browser execution.
How should an agent handle conflicting sources?
Do not choose the most convenient source or average incompatible claims. Record each claim, its evidence, date, scope, and source identity. Then search for the reason for the difference: version, geography, methodology, event time, or an upstream correction. If the conflict remains, narrow the conclusion and expose both supported conditions. A report can be decision-useful while preserving uncertainty, but it should not manufacture consensus that the evidence does not provide.
How many sources should a research agent collect?
Set a coverage rule instead of a fixed source target. A simple factual lookup may need one authoritative primary source, while a disputed decision may require independent corroboration and explicit counterevidence. Stop when required subquestions and material claims satisfy the contract, not when the agent reaches an arbitrary count. Also cap queries, tool calls, time, and credits so a low-value gap does not create an endless search loop.
Can a deep research agent replace human review?
It can reduce collection and traceability work, but the review boundary depends on consequence. Human review remains appropriate when sources conflict, access or licensing is unclear, private data is involved, or the output affects legal, medical, financial, safety, or other high-impact decisions. The agent should make that review easier by exposing claim-level evidence, failures, and gaps. It should not use fluent language or a confidence score as a substitute for accountable judgment.






