How to Keep Duplicate Sources Out of a Research Agent
Group repeated text without losing original reporting, corrections, or provenance. Use conservative URL keys, validated snapshots, and a reviewable source ledger.
AnyCrawler guides
Learn how to find web sources, extract readable content, and handle JavaScript pages in AI workflows. Browse guides by topic and choose the next task you need to solve.
Published guides
14 guides
Group repeated text without losing original reporting, corrections, or provenance. Use conservative URL keys, validated snapshots, and a reviewable source ledger.
Keep useful documents when individual reads fail. Classify errors, bound retries and timeouts, validate content, and preserve an auditable partial result.
Preserve requested, final, and canonical URL evidence, verify the cited passage before replacing its address, and handle redirects and missing API identity fields explicitly.
A decision framework for AI agents that pairs machine-readable page content with screenshot evidence when layout, visible state, or human review changes the meaning of a claim.
Google retired its official News Search API. Compare RSS, third-party SERP services, dedicated news APIs, and an evidence-first monitoring workflow.