Keep duplicate sources out of a research agent by separating URL aliases, repeated document text, and reports about the same event. Normalize addresses conservatively, compare validated content snapshots, and treat canonical declarations or similarity scores as evidence to investigate. Preserve every observation in a source ledger even when you send only one copy of its text to the model. This prevents syndicated material from crowding out independent reporting without deleting distinct pages from the same publisher, corrections, or competing accounts. The practical goal is less repeated context with recoverable provenance, not the smallest possible source list.

Decide what you are trying to merge

A search result is an observation from a query. A fetched page is a document snapshot. A news event can have several documents describing it. Giving all three the same deduplication key creates avoidable mistakes: a corrected article may disappear because its address stayed the same, while a copied press release may look like independent confirmation because it appears on another domain.

Use separate relationships and actions:

Observation Decision What must survive
Same request under the same collection conditions Reuse a read only within the task's freshness policy All queries and discovery positions
Different URLs, identical validated extracted text Share a context body; retain separate observations URLs, capture times, publisher identities
Declared canonical points to another page Review the target before joining document identities Raw declaration, verification result, original snapshot
Similar text with a changed number or negation Keep both until the difference is explained Claim-changing passage and version relationship
Same event, different reporting or interviews Group under an event; keep distinct documents Independent evidence and disagreements
Same publisher, unrelated original pages Keep both Each document's task relevance
Empty page, access message, or incomplete extraction Mark content invalid; exclude from evidence grouping Failure state and recovery decision

Sharing a context body is reversible: the model reads the repeated text once, while the ledger still explains every place it was observed. Deleting the source record makes that reconstruction much harder.

Normalize URLs without erasing page meaning

Use a URL parser rather than lowercasing an entire string. Scheme and host case can be normalized; path case and query values may carry meaning. The query component participates in resource identity under the URI syntax and comparison rules. Removing every parameter can collapse distinct languages, editions, filters, or pagination into one address.

Start with a conservative key. Keep parameter order, repeated parameters, path case, and fragments unless a verified rule for that source permits a transformation. Ordinary section fragments can often share a document read, but fragment-driven application routes may select different content. Preserve the submitted URL either way. HTTP and HTTPS, trailing-slash variants, or different hosts should remain distinct until observed behavior or a reviewed source policy establishes equivalence.

For tracking parameters, compare the original and cleaned representations before adding a source-specific removal rule. A successful comparison applies to the tested representation and collection conditions; it does not authorize rewriting every website's URLs. Keep a policy version alongside the key so that a later rule change does not silently rewrite old evidence.

URL grouping also needs a time boundary. Reusing a request within one research run can be appropriate, while suppressing that URL forever would hide corrections and new editions. Deduplication does not choose a freshness policy for you.

A canonical declaration creates a candidate relationship

Retain requested, final, and declared canonical URLs as separate fields. Before merging on a canonical value, retrieve its target when permitted and compare document identity, relevant passages, language, and version. A declaration that drops an important query parameter or points to a generic overview may erase the evidence your agent needs.

The canonical-versus-final-URL verification workflow walks through that check, including cases where a reachable destination no longer supports the claim. Use it before substituting a canonical address in the ledger. If verification is unavailable or the content disagrees, retain an unresolved relationship instead of assigning the documents one identity.

Canonical validation and exact-text sharing answer different questions. Identical validated text can share a context payload even while publisher attribution remains unresolved. A verified common publication identity can still contain several snapshots with meaningful edits. Keep those dimensions separate.

Start with a conservative exact-snapshot implementation

The following Node.js example leaves every input row in a ledger and shares text only after the caller has validated its extraction. It hashes the exact Markdown bytes, includes an extraction-version key, and checks text equality after a hash match. It deliberately leaves canonical merging and near-duplicate decisions out of automatic grouping.

Save it as dedup.mjs and run node dedup.mjs in an environment supporting ES modules. The fixture statements are invented test data, not reports about a real bridge. The assertions check a repeated body, a distinct same-domain document, a changed snapshot at the same URL, an invalid empty response, and URL distinctions that must survive.

import { createHash } from 'node:crypto';
import assert from 'node:assert/strict';

const digest = text => createHash('sha256').update(text, 'utf8').digest('hex');

export function urlKey(raw) {
  const u = new URL(raw);
  if (!['http:', 'https:'].includes(u.protocol) || u.username || u.password) {
    throw new Error('Expected a public HTTP(S) URL without credentials');
  }
  return u.href; // Keep query order, parameters, path case, and fragment.
}

export function groupSnapshots(rows) {
  const ledger = [];
  const groups = [];
  const byHash = new Map();
  for (const row of rows) {
    const saved = { ...row, action: 'review', reason: 'content_invalid' };
    ledger.push(saved);
    if (!row.contentValid || !row.markdown?.trim()) continue;
    try {
      saved.urlKey = urlKey(row.requestedUrl);
    } catch {
      saved.reason = 'url_invalid';
      continue;
    }
    saved.contentHash = digest(row.markdown);
    // Include the extraction contract so different readers are not conflated.
    const bucketKey = JSON.stringify([row.extractorVersion, saved.contentHash]);
    const candidates = byHash.get(bucketKey) ?? [];
    const match = candidates.find(g => g.markdown === row.markdown);
    if (match) {
      match.members.push(row.id);
      saved.groupId = match.id;
      saved.action = 'reuse_context';
      saved.reason = 'exact_validated_snapshot';
    } else {
      const group = { id: row.id, markdown: row.markdown, members: [row.id] };
      groups.push(group);
      candidates.push(group);
      byHash.set(bucketKey, candidates);
      saved.groupId = group.id;
      saved.action = 'keep_context';
      saved.reason = 'distinct_snapshot';
    }
  }
  return { ledger, groups };
}

const row = (id, path, markdown, extra = {}) => ({
  id, requestedUrl: `https://example.com/${path}`,
  extractorVersion: 'demo-markdown-v1', contentValid: true, markdown, ...extra
});
const result = groupSnapshots([
  row('original', 'report', 'The bridge is open.'),
  row('alias', 'report?utm_source=demo', 'The bridge is open.'),
  row('other', 'interview', 'The repair crew describes a different route.'),
  row('correction', 'report', 'The bridge is not open.'),
  row('empty', 'empty', '', { contentValid: false })
]);
assert.equal(result.groups.length, 3);
assert.equal(result.ledger.length, 5);
assert.deepEqual(result.groups[0].members, ['original', 'alias']);
assert.equal(result.ledger[4].action, 'review');
assert.notEqual(urlKey('https://example.com/?page=1'), urlKey('https://example.com/?page=2'));
assert.notEqual(urlKey('https://example.com/A'), urlKey('https://example.com/a'));
assert.notEqual(urlKey('https://example.com/#/a'), urlKey('https://example.com/#/b'));
assert.equal(urlKey('https://EXAMPLE.com:443/'), 'https://example.com/');
console.log(JSON.stringify(result, null, 2));

This controlled run produces three context groups from five ledger rows. The original and alias share a body; the interview and correction remain separate; the empty response remains in review. Those are fixture outcomes, not a measured deduplication accuracy or token-saving benchmark. The first encountered member supplies an internal group identifier; it is not automatically the original publisher or the preferred citation.

contentValid is an upstream contract, not a check that a string is long enough. Confirm that the retrieved material contains the expected document, relevant passage, and necessary qualifications. A common access-denied template should fail validation even if its hash matches many responses. The example's final nonempty check merely adds a defensive guard.

Check the adapter against an actual public response

Use AnyCrawler's free URL-to-Markdown tool to inspect a permitted public page before connecting a reader to the grouping function. Its public demonstration endpoint can return cached results. Search discovers candidates; Fetch reads available HTML, Render addresses content that needs JavaScript, and Screenshot preserves visible context. None of those steps establishes source independence, and rendering does not imply interactive browser automation.

This probe reads the same documentation-example page with and without a campaign parameter. It records the outer API status separately from the target status and leaves unavailable provenance fields as null.

import { createHash } from 'node:crypto';
import { writeFile } from 'node:fs/promises';

const observations = [];
for (const target of ['https://example.com/', 'https://example.com/?utm_source=dedup-demo']) {
  const endpoint = new URL('https://api.anycrawler.com/free/v1/crawl');
  endpoint.searchParams.set('url', target);
  const response = await fetch(endpoint, { signal: AbortSignal.timeout(30000) });
  const body = await response.json();
  const markdown = body.results?.markdown;
  const valid = response.ok && body.status_code === 200 &&
    typeof markdown === 'string' && markdown.includes('Example Domain');
  observations.push({
    requestedUrl: target, observedAt: new Date().toISOString(),
    gatewayStatus: response.status, targetStatus: body.status_code ?? null,
    finalUrl: body.final_url ?? body.results?.final_url ?? null,
    canonicalUrl: body.canonical_url ?? body.results?.canonical_url ?? null,
    creditsUsed: body.credits_used ?? null,
    contentValid: valid, markdown: markdown ?? null,
    contentHash: valid ? createHash('sha256').update(markdown).digest('hex') : null,
    response: body
  });
}
await writeFile(new URL('live-example.json', import.meta.url), JSON.stringify(observations, null, 2));
console.log(JSON.stringify(observations, null, 2));
if (observations.some(x => !x.contentValid)) process.exitCode = 1;

Both tested requests returned gateway and target status 200, the title “Example Domain,” and identical Markdown. Their exact UTF-8 SHA-256 was 5945db6fd8137aa377638814ca9bb1ac0a663fd90a97f11f86c3f5c09cfb40e3. The response did not expose a target final URL, canonical URL, or credits-used value. Missing credits are not zero credits, and the API endpoint's own response URL is not the target's final URL.

This small probe demonstrates equality for those two observations. It does not prove that every campaign parameter is harmless, that every page can be read, or that the response was freshly collected. A production adapter should preserve its actual reader configuration, cache information where exposed, and the extraction contract used for comparison. Apply explicit fetch or render selection where needed; do not invent missing response fields.

Use near-duplicate scores to find review candidates

Exact hashes miss lightly edited copies. One way to propose pairs is to divide text into overlapping sequences of words, called shingles, and compare the sets with Jaccard overlap: intersection size divided by union size. The Stanford information-retrieval explanation of shingling describes this approach. Similarity provides a candidate signal; it does not determine whether a difference matters to the research question.

Choose tokenization, shingle length, and a review threshold using labeled examples from your own documents. Include copied articles with added navigation, short notices, translations, corrections, and related reports with different evidence. A high score can hide a changed deadline or the word “not”; a low score can occur between translations of the same underlying reporting. Do not adopt a universal score as a deletion rule.

During review, compare the passages that support the claim. Record whether a pair is a reproduction, an updated version, a partial quotation, or simply related. Shared headlines and embedding similarity can help find candidates, but neither should erase documents by itself. Avoid blindly joining chains of similar pairs: a document related to an intermediate copy may still differ materially from the group's representative.

Track editorial independence after text grouping

Different domains can publish the same underlying material. Research on news provenance and text reuse examines how republished content spreads. In an agent, the practical consequence is to record attribution and reporting lineage instead of treating domain count as a count of independent confirmations.

For each retained document, keep the requested and observed URLs, publisher, byline when available, publication and collection times, snapshot hash, relevant passage, and access state. Add the context group, document-version relationship, event identifier if established, decision reason, and the evidence supporting that relationship. Unknown fields should remain unknown.

A proposed ledger entry might say: “share the body with observation A; retain publisher B and its attribution; independent reporting unresolved.” Another might say: “same event, separate interview, keep for a claim absent from A.” These records make source selection explainable without suggesting that the agent has established authorship merely by comparing text.

Apply diversity during selection after resolving clear repetitions. Prefer documents that add relevant evidence, primary material, a distinct perspective, or a substantive disagreement. Keep an additional page from the same publisher when it answers a different part of the question. If the shortlist is still redundant, search for the specific missing evidence rather than filling an arbitrary domain quota.

How much repetition belongs in the final context?

The remaining tradeoff is between a compact model input and the provenance a reviewer may need. A task about a published statement might use one body plus its observation ledger. A task about how that statement changed or spread may need several versions and republications as the subject of the research itself.

Choose the merge policy from that distinction. Before enabling automatic near-duplicate suppression, inspect cases where a rejected document would have changed the answer. Preserve the decision and an undo path so that the policy can become more conservative without recollecting lost evidence. The useful question is whether your compressed context still exposes the differences that would make the agent revise its conclusion.

Frequently asked questions

Should a research agent deduplicate URLs or page content first?

Use conservative URL keys before collection to avoid redundant requests under the same conditions, then compare validated content snapshots after collection. These stages solve different problems. Different addresses can return the same text, while one address can change over time. Keep original discovery records and use a separate freshness policy so that request reuse does not suppress corrections or later editions.

Is a matching content hash enough to merge sources?

A matching hash can identify candidate equal snapshots, but the text must be valid evidence and the extraction contract must match. The example also compares the actual text before sharing a context body. Preserve both observation records, including publishers and timestamps. Equality does not establish which publisher originated the material, whether the statement is true, or whether the two appearances are independent confirmation.

Can I remove all query parameters before deduplication?

No. Parameters may select a language, document version, page, filter, or other meaningful state. Start by retaining them, including their order, and introduce a source-specific cleanup rule only after checking equivalent representations. Keep the original address and policy version in the ledger. The public example supports equality for its tested campaign variant; it does not establish a general rule for other parameters or websites.

Should reports about the same event become one source?

Treat the event as a grouping relationship while preserving distinct documents. Several reports may contribute different interviews, observations, or contradictory evidence even when their headlines resemble each other. Reproductions can share context when verified, but a common event is not a sufficient merge reason. Select material according to the claims it supports and retain uncertainty about reporting lineage when attribution cannot be established.

What similarity threshold should I use for near duplicates?

Set a candidate-review threshold from labeled examples that resemble your task, then inspect false merges as well as missed copies. There is no universal value justified by this workflow. Corrections, short notices, translations, and common templates need special attention because overlap can misrepresent their significance. Keep consequential differences in the context even when most of a document matches another, and record the review reason.

Does one source per domain guarantee a diverse answer?

No. Different domains can repeat the same underlying reporting, and one publisher can host several original documents that answer different parts of a question. Use domain variety as a selection consideration after checking text reuse and relevant evidence. Preserve attribution, bylines when available, and unresolved lineage. A useful source set adds independent support or meaningful disagreement instead of merely satisfying a domain count.