An agent should cite a URL that readers can open and use to verify the specific claim it extracted. Keep the requested URL, the final URL after redirects, and the publisher’s declared canonical URL as separate fields. Prefer a canonical citation only after checking that its destination preserves the relevant content and publication identity; otherwise retain the verified final page or flag the citation for review. A redirect records where a request went, while a canonical declaration proposes which version represents the content. Neither replaces checking the evidence itself.

Give each URL a different responsibility

The requested URL explains how the pipeline entered a source: a search result, a saved bookmark, or a link containing campaign parameters. It belongs in the collection record even when it is unsuitable as the reader-facing citation.

The final URL identifies the response actually reached by that request. For a direct web request, the Fetch API’s Response.url property reports the URL after redirects. When calling a scraping service, the HTTP response URL is the service endpoint; use the service’s documented target-page field for the source destination. Confusing those layers can produce a citation to an API instead of an article.

The canonical URL is a declaration of preferred publication identity. The canonical link relation described in RFC 6596 concerns duplicate content or a representation containing the referring page’s content. It can be self-referential, relative, or cross-domain, and can appear in HTML or an HTTP Link header. It is a claim to evaluate rather than proof that the destination is suitable for your quotation.

Google’s selected canonical is another concept. Google considers multiple signals and treats a publisher’s canonical preference as a hint. Reading a page’s tag does not reveal Google’s indexing decision. Your agent can select a defensible citation without trying to reproduce search indexing.

Make the citation decision at the claim level

A page can have a plausible canonical destination that omits the exact detail you need. Consider a regional product page whose canonical points to a global product overview: the overview may identify the product but omit the regional price. The decision must follow the statement being cited, not just the matching product name.

Use this application policy as a starting point. The examples are decision scenarios, not claims about a particular publisher’s configuration.

Observed situation Check before selecting a citation Suggested outcome
Final page declares itself canonical Confirm the page contains the cited passage Cite that page; retain all original fields
Tracking URL declares a clean canonical Read the clean destination and compare the relevant passage Use the clean URL if identity and evidence agree
Request redirects to another domain Check the destination’s publisher, page identity, and content Cite the verified destination; preserve the entry URL
Canonical points to a homepage or unrelated article Compare the claim with the proposed destination Reject the canonical substitution; keep the evidence page
Canonical target is unavailable or redirects repeatedly Record its outcome and enforce a bounded resolution policy Leave canonical verification unresolved
No canonical or final URL is exposed by the reader Inspect the reader’s contract; obtain missing evidence separately Keep missing fields null; do not infer them from the title
Several conflicting canonical declarations appear Preserve the declarations and their locations Require review before choosing a preferred identity

Tracking parameters deserve particular care. A parameter name that looks disposable is not proof of equivalent content. Removing a query may change the selected language, edition, currency, filters, or document version. Preserve the original address, then verify a proposed clean address against the claim. Do not apply a global “delete everything after the question mark” rule.

Capture URL evidence without filling gaps with guesses

The following Python probe runs against AnyCrawler’s public no-key endpoint using a permitted test page. It intentionally distinguishes an API request succeeding from a citation becoming ready. Save it as a Python file and run it with Python and curl installed.

import json
import subprocess
from urllib.parse import urlencode

requested = "https://anycrawler.com/crawler/page/fetch/"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urlencode(
    {"url": requested}
)
reply = subprocess.run(
    ["curl", "--silent", "--show-error", "--fail",
     "--max-time", "30", endpoint],
    check=True, capture_output=True, text=True, encoding="utf-8",
)
data = json.loads(reply.stdout)
body = data.get("results", {}).get("markdown") or ""
record = {
    "requested_url": requested,
    "reported_requested_url": data.get("requested_url"),
    "final_url": data.get("final_url"),
    "canonical_url": data.get("canonical_url"),
    "target_status": data.get("status_code"),
    "title": data.get("results", {}).get("title"),
    "credits_used": data.get("credits_used"),
    "has_content": bool(body.strip()),
    "citation_url": None,
    "citation_state": "NEEDS_VERIFICATION",
}
print(json.dumps(record, indent=2))

In the publication run, the endpoint returned target status 200 and the Fetch page’s title and Markdown. It did not expose final URL, canonical URL, or credits used. Those fields therefore remained null. Missing credits are not evidence of zero credits, and a successful target status does not establish publication identity. The public endpoint may serve cached content; this probe is a response-contract check, not an authenticated Fetch benchmark.

For the production integration, open the AnyCrawler Fetch page API’s request and response contract, then map its documented requested, final, and canonical URL fields into separate entries. Start with a permitted URL you know, inspect the actual response, and only advance the record after verifying the passage you intend to cite. This is the next step when the no-key probe does not expose enough provenance.

Search helps discover candidate pages. Fetch reads a known page; Render is appropriate when the required content depends on browser execution. A screenshot can preserve visible context, but it does not independently establish a canonical relationship. Interactive login, clicks, and form submissions belong to a separate browser automation workflow. Changing the read path should create another observation with its own fields, not overwrite the previous observation’s identity.

Reproduce the cases that break simple URL selection

Separate live source observations from controlled policy tests. A public page can change later, while a deliberately incorrect fixture provides a stable way to check rejection behavior without alleging that a real publisher has broken metadata.

The live checks in this publication run produced these results:

  • The Fetch documentation URL returned 200 and declared that same address as canonical.
  • Adding ?utm_source=citation-test preserved the query in a direct request’s final URL, while the HTML declared the clean documentation address. The no-key extraction returned identical Markdown for the clean and parameterized inputs in this run.
  • A public httpbin redirect fixture pointing to https://example.com/ reached that cross-domain destination in a separate direct HTTP request. The no-key service returned the Example Domain content but omitted final URL, demonstrating why content success alone cannot repair a missing identity field.

The direct requests and the extraction calls are separate observations. Their results should not be fused into a claim that both clients saw the same redirect chain or representation. Store the client, time, status, and content associated with each observation. If the application must prove the redirect history, use a reader that exposes each hop and retain that history.

Test a deliberately wrong canonical with an isolated policy fixture:

def choose_citation(final_url, canonical_url, *, final_verified,
                    canonical_verified, claim_present):
    if not final_verified:
        return None, "REVIEW"
    if canonical_url and canonical_verified and claim_present:
        return canonical_url, "VERIFIED_CANONICAL"
    return final_url, "VERIFIED_FINAL"

# Synthetic inputs: the declared homepage does not contain the claim.
assert choose_citation(
    "https://example.com/article", "https://example.com/",
    final_verified=True, canonical_verified=True, claim_present=False,
) == ("https://example.com/article", "VERIFIED_FINAL")

This function consumes verification outcomes; it does not perform verification. A production caller must establish publisher and document identity as well as passage presence before setting canonical_verified. The fixture checks the decision boundary when a reachable canonical destination loses the claim. Add cases for an inaccessible target, a different language, a cross-domain publisher change, and a missing final URL according to your application’s needs.

Keep the evidence record separate from the display link

Store the collection inputs and observations as a durable record, then attach the chosen citation and its reasoning. This lets you replace a broken display link without rewriting the history of what the agent actually read.

Record field Purpose
requested_url and discovery context Reproduce the entry point and explain source selection
final_url, status, and redirect observations Identify the reached representation and transport outcome
canonical_raw, resolved candidate, declaration location Preserve the publisher’s claim and how it was interpreted
reader, capture time, cache information Describe the observation’s execution and freshness boundaries
passage, section, and content hash Connect the statement to retained evidence
canonical_verification and reason Explain acceptance, rejection, or missing evidence
citation_url and decision time Record the address presented to the reader

If the canonical value is relative, resolve it using the applicable document or header context and preserve the raw value. Do not treat HTML parsing and HTTP Link header parsing as interchangeable string searches. Retain conflicting values rather than silently accepting whichever a parser returns first.

A matching content hash can support an exact-representation comparison. Different hashes require interpretation: navigation, personalized content, or small edits can change bytes without changing the cited passage. For a quotation, check the quotation and its surrounding qualification. For a numerical claim, check the entity, unit, applicable period, and conditions. These are proposed application checks, not guarantees supplied by a canonical tag.

How much canonical verification is enough?

The remaining design choice is how much additional retrieval a citation deserves. Re-reading every canonical target adds work, but automatically replacing every final URL can erase useful evidence. An application could prioritize verification for cross-domain changes, disappearing query parameters, or claims sensitive to edition and locale, while leaving unresolved citations explicitly marked for review.

The threshold should follow the consequence of a wrong link. A reading list can tolerate a page-level citation that a human checks later; an automated report quoting a specific policy exception needs the destination to retain that exception and its context. Decide what evidence must survive before optimizing deduplication. The useful open question is whether your citation policy preserves the claim when a publisher’s preferred URL changes tomorrow.

Frequently asked questions

Should an AI agent always cite the canonical URL?

Use a canonical address when verification establishes that it preserves the cited content and the relevant publication identity. A declared preference can lead to an unrelated overview, a different edition, or an unavailable page. Keep the final page as the evidence location while checking the candidate. If substitution fails, retain a verified final-page citation or require review instead of presenting the canonical address as automatically authoritative.

What is the difference between requested URL and final URL?

The requested URL is the address submitted to the reader, while the final URL is the address reached after that reader follows redirects. Keep both because the original input explains discovery and the destination identifies the response. When using a scraping API, distinguish its endpoint URL from the target page’s final URL. If the API does not return the latter, record the gap explicitly rather than inventing a fallback.

Can a canonical URL point to another domain?

A cross-domain canonical is possible, but your citation policy still needs to validate the destination. Read the target, check who published it, and confirm that the passage supports the same statement with the same qualifications. Preserve the observed evidence page even if you accept the preferred identity. If access, content, or publisher identity remains uncertain, keep the canonical candidate unresolved and do not silently replace the citation.

Is it safe to remove tracking parameters before citing a page?

Treat a cleaned URL as a candidate until you verify equivalent evidence. Some parameters change only attribution, while others select language, version, filters, or other meaningful state. Preserve the original request and compare the clean page with the passage you used. The tested documentation example supported removing its campaign parameter, but that result does not establish a safe rule for every parameter or every website.

Does a self-canonical tag prove that the page is the original source?

No. It tells you that the page declares itself preferred; it does not establish authorship, originality, accuracy, or the publication history of its claims. Check the publisher and the supporting passage separately. Keep a self-canonical page as a candidate citation when it satisfies the task, but do not let the tag substitute for evaluating whether the page actually supports the statement your agent will make.

What should happen when canonical and final URL disagree?

Preserve the disagreement and investigate why it matters for the claim. A parameterized duplicate may justify a clean citation, while a redirect to a different edition or an unrelated canonical target may not. Save both observations and the verification result. Choose the address that lets readers recover the evidence, and leave a review state when neither destination can be verified sufficiently for the application’s purpose.