Extract links from a web page by treating every link as a record, not a string. Save the source page, raw href, resolved destination, anchor text, nearby sentence or block, and an internal/external label. Start from the page version your workflow actually needs, then verify that expected links are present before enqueueing anything. This produces a reviewable one-page link set. It does not discover every URL on a domain, prove that destinations work, or justify stripping query strings and fragments without checking what they mean.
A useful link record answers why the URL was present
A bare destination such as https://example.com/docs tells a later agent where it could go. It does not say which page supplied the link, what the anchor promised, or whether the link appeared in a navigation menu, a citation, a product card, or a legal notice.
The minimum record depends on the downstream task, but this shape is a practical starting point:
| Field | Why keep it | What it does not prove |
|---|---|---|
sourceUrl |
Identifies the page where the link was observed | That the source is authoritative |
rawHref |
Preserves the page author's exact value | That the value is a valid web URL |
resolvedUrl |
Gives the absolute destination after base resolution | That the destination is reachable |
anchor |
Records the clickable label or image alternative | That the destination matches the label |
context |
Keeps the nearby sentence, heading, list item, or card | That the surrounding claim is correct |
internal |
Separates same-origin from external destinations | That either class should be followed automatically |
observedAt |
Ties the record to an extraction run | That the link remained unchanged later |
This record also preserves evidence when filtering changes. A research agent may want only cited sources, while an internal audit may want navigation links. Keeping the original context lets each consumer make that choice without guessing from a URL string.
Google's current guidance on crawlable links and anchor context gives a concrete reason for retaining more than href: a normal anchor needs an href, and descriptive anchor text plus its surrounding sentence helps people and crawlers understand the destination. Link extraction should keep those signals together even when the final use is not SEO.
Choose the page version before selecting anchors
The first question is whether the required links exist in the HTML response or appear only after JavaScript runs. Fetching the wrong page version can produce a clean, valid-looking list that omits the links the task actually needs.
For a known public URL whose HTML already contains the relevant body, the AnyCrawler Fetch Web Page API documents the direct path: send method: "fetch" and request include_links: true when structured link results are needed. The page also exposes requested, final, and canonical URL fields when available, which should travel with the extracted records. If a task-specific link is absent because a public page builds that section in JavaScript, Render may be the appropriate escalation. JavaScript somewhere on the page is not enough reason by itself.
Define a content contract before escalating. A documentation job might require a link whose anchor contains “API reference.” A research job might require citations inside the main article. If Fetch returns those links, rendering adds cost and another source of variation without closing an evidence gap. If it does not, inspect whether the missing section is browser-built before trying Render once with a clear wait condition.
The scope remains one page in either mode. Following every extracted destination, discovering a site's unlinked pages, reading a sitemap, and respecting crawl budgets are separate jobs.
Select anchors, resolve destinations, and keep raw identity
When the page is available as a DOM, select elements that actually represent the link contract. The browser's document selector API returns matching elements in document order, so document.querySelectorAll("a[href]") is a sound starting point for normal HTML anchors. Narrow the selector to the main article or another task-owned container when menus and footers would overwhelm the record set.
Resolution comes after selection. The URL constructor with a base URL handles root-relative paths, directory-relative paths, parent segments, and protocol-relative references according to URL rules. String concatenation does not. Keep invalid or unsupported schemes as structured rejections rather than silently turning them into destinations.
A safe transformation keeps both sides:
- Store the exact
hrefattribute asrawHref. - Resolve it against the effective base URL, normally the fetched page's final URL unless a trusted page base changes that rule.
- Store the resolved URL separately.
- Classify the resolved origin against the source origin.
- Capture a bounded context unit such as the nearest sentence, list item, table row, card, or heading plus paragraph.
- Record exclusions such as empty anchors,
mailto:,tel:,javascript:, and malformed values with a reason.
Do not normalize away evidence during extraction. A fragment may identify the cited section. A query parameter may select a language, product, version, or document state. Tracking parameters can be candidates for a later canonicalization policy, but the original observed value should remain available for review.
Reproduce a one-page extraction before trusting the pipeline
The following Node.js example uses AnyCrawler's public free endpoint and the permitted Example Domain fixture. The free response exposes Markdown rather than the authenticated results.links option, so the example extracts Markdown links and then applies the same record principles. The parser is deliberately narrow for this controlled fixture; production Markdown with nested labels or parenthesized destinations should use a real Markdown parser.
import { createHash } from "node:crypto";
const sourceUrl = "https://example.com/";
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", sourceUrl);
const response = await fetch(endpoint, {
headers: { accept: "application/json" },
signal: AbortSignal.timeout(30_000),
});
if (!response.ok) throw new Error(`AnyCrawler HTTP ${response.status}`);
const data = await response.json();
const markdown = data.results?.markdown;
if (data.ok !== true || data.status_code !== 200 || typeof markdown !== "string") {
throw new Error("The page response failed its content contract");
}
const links = [...markdown.matchAll(/\[([^\]]+)\]\((https?:\/\/[^)\s]+)\)/g)].map(
(match) => {
const resolved = new URL(match[2], data.final_url ?? data.requested_url ?? sourceUrl);
const start = Math.max(0, match.index - 70);
const end = Math.min(markdown.length, match.index + match[0].length + 70);
return {
sourceUrl: data.requested_url ?? sourceUrl,
anchor: match[1].replace(/\s+/g, " ").trim(),
rawHref: match[2],
resolvedUrl: resolved.href,
internal: resolved.origin === new URL(sourceUrl).origin,
context: markdown.slice(start, end).replace(/\s+/g, " ").trim(),
};
},
);
if (links.length < 1) throw new Error("No Markdown link survived extraction");
console.log({
apiStatus: response.status,
pageStatus: data.status_code,
title: data.results.title,
links,
markdownSha256: createHash("sha256").update(markdown, "utf8").digest("hex"),
creditsUsed: data.credits_used ?? null,
});
The September 7, 2026 run returned API HTTP 200, page status 200, the title Example Domain, and one record for the Learn more link. The Markdown SHA-256 was 5945db6fd8137aa377638814ca9bb1ac0a663fd90a97f11f86c3f5c09cfb40e3. Final URL, canonical URL, and credits were not exposed, so the ledger kept them as null. This proves the fixture's one-page path on that run. It does not establish link coverage for other pages, authenticated include_links behavior, or destination health.
Preserve repeated destinations when their contexts differ
Deduplication is usually where a useful record collapses back into a bare URL. A page may link to the same destination from the navigation, a comparison table, and a cited sentence. Those are one destination but three observations with different meaning.
Model the relationship in two layers. A destination table can hold the resolved URL and later checks such as status or canonical identity. An observation table can hold source page, anchor, context, DOM region, and observed time. This allows queueing code to fetch a destination once while review code retains every reason it appeared.
Apply filtering after the observation exists. Common policies include ignoring same-page fragments for a crawl queue, excluding sign-in and account routes from a public research run, or keeping only main-content links for citation review. Store the policy name and rejection reason. A later operator can then distinguish “the page had no links” from “the extractor found links and the current policy excluded them.”
Context should also have a stable boundary. A full page is too noisy; a fixed character window can split a sentence or mix neighboring cards. Prefer the smallest semantic container that explains the link: the containing sentence, list item, table row, figure caption, or card. Add the nearest heading when the local block would otherwise be ambiguous.
Classify failures before retrying or following anything
Link extraction success has several layers. A successful page request only clears the first one.
| Observed state | Classification | Next action |
|---|---|---|
| Request failed with a temporary transport or upstream condition | Retrieval failure | Retry only when marked temporary, with a finite limit |
| HTTP status is successful but expected body or anchors are absent | Content-invalid | Check for an empty shell, wrong page, consent view, or Render requirement |
Anchors exist but href values are empty, malformed, or unsupported |
Link-invalid | Preserve rejected observations and reasons; do not enqueue them |
| URLs resolve but anchors and context are missing | Record-incomplete | Re-extract from the DOM or Markdown with a context-aware contract |
| One destination appears many times | Duplicate destination | Fetch once if appropriate; keep distinct observations |
| Destination later redirects, fails, or requires authorization | Follow-up failure | Store that result separately from the original observation |
| Link set contains only navigation when citations were required | Task-invalid | Narrow the extraction region or return an explicit evidence gap |
Verification should test the task, not the row count. Confirm that known anchors are present, relative references resolve against the correct base, internal/external labels use the intended origin rule, excluded schemes have reasons, and sampled contexts explain why each link appeared. If the output will drive requests, apply destination allowlists, access rules, concurrency limits, and URL safety checks in that later stage.
How much context should a link record retain?
There is no universal context window. A citation reviewer may need the whole sentence and nearest heading. A navigation audit may need the menu or footer region. A data pipeline may need a table row with column labels. Storing too little makes the record ambiguous; storing too much repeats the page and raises storage and review cost.
The open design question is which context unit predicts a correct downstream decision for a specific workflow. Measure that decision directly. Track how often reviewers can classify a link without reopening the source, how often a context boundary mixes unrelated claims, and whether a heading changes the interpretation. Keep the raw observation addressable so the policy can evolve without pretending that today's context rule is part of the URL itself.
Frequently asked questions
Is extracting links from one page the same as crawling a website?
No. One-page extraction reports links observed in a particular fetched or rendered document. A site crawl follows eligible destinations across multiple pages, manages duplicate URLs and crawl limits, and may use sitemaps or other discovery sources. Label the one-page result with its source URL and observation time. If the next stage follows links, describe that stage separately and apply its own scope, access, retry, and stop rules.
Should duplicate URLs be removed immediately?
Keep duplicate observations until their context has been recorded. The same resolved destination may appear in navigation, a comparison table, and a citation sentence. A fetch queue can later collapse those observations to one destination request, while an audit still retains each anchor and reason the link appeared. Immediate URL-only deduplication saves rows but can erase the evidence needed to rank, filter, or explain the link.
Why store both the raw href and the resolved URL?
The raw href preserves what the source page actually contained, including relative paths, fragments, and query strings. The resolved URL gives downstream code an absolute destination based on the effective page URL. Keeping both makes normalization reviewable and lets you correct a bad base rule later. It also prevents a canonicalization policy from becoming indistinguishable from source evidence. Neither value alone proves that the destination is reachable.
When should link extraction use Render instead of Fetch?
Use Render when the links required by the task appear only after JavaScript builds the public page state. Start with Fetch when the HTML already contains the relevant anchors, then check for expected links or sections. Do not escalate just because the site uses JavaScript somewhere. If Render is necessary, set a clear wait condition and validate the resulting link set again; browser execution does not turn a one-page read into interaction or site-wide discovery.
Can I strip fragments and query parameters before deduplication?
Only under a policy that understands their meaning for the target site and task. Fragments can identify a cited section, while query parameters can select language, version, filters, or state. Keep the observed and resolved forms first. A later queue may derive a fetch key that removes known tracking parameters or same-page fragments, but it should record that transformation and preserve the original observation for review.
How do I verify that extracted links are usable?
Validate in layers. Confirm the correct source page and page status, require known anchors or regions, resolve sample relative paths against the effective base, and inspect context for a few records. Check that exclusions have reasons and that internal/external classification follows the intended origin rule. Destination requests are a separate verification step: run them only within scope, then store redirects, failures, canonical hints, and access boundaries without rewriting the original link observation.






