To chunk web pages for RAG without losing citations, keep each page's source URL and title before splitting, then use trustworthy headings or other document boundaries to form the first chunks. Attach the heading path and a stable order to every chunk. If a section is still too large for your embedding or retrieval setup, split inside that section while carrying the same source metadata forward. A retrieved passage should lead back to the page and the exact section that supports the answer; a clean text fragment alone cannot do that.

Preserve the page before you split it

Chunking starts after access and extraction have succeeded. A page may return a successful HTTP status while its Markdown contains only navigation, or it may lose a table footnote that changes the meaning of a row. The page-to-Markdown quality checks cover those upstream failures. This workflow begins once the required body and section markers are present.

Store the requested URL, any final or canonical URL actually returned, the page title, retrieval time, and the Markdown snapshot. Do not fill missing URL fields by guessing: a requested URL is evidence of what you asked for, while a final URL or canonical URL answers a different identity question. Keep the original page snapshot so a later citation can be checked against what the system actually read.

Microsoft's chunking guidance for RAG describes both fixed-size and structure-based methods. HTML and Markdown headings can define sections, but one page can still need further splitting when a section exceeds the limits of the downstream model. There is no useful universal chunk size without the model, document shape, query pattern, and evaluation set.

Define what a citable chunk carries

Field Keep it because Fail the chunk when
source_url The reader needs a retrievable source page. It is absent or points to a different document.
title and heading_path The passage needs its surrounding topic. The heading path describes another section.
ordinal and exact text Reviewers need the passage's position and wording. A split cuts a required table row, code block, or qualification.
Retrieval time and snapshot ID A later page version may differ. The system cannot say which version supplied the answer.

This is an application-level record, not a claim that AnyCrawler returns an indexed RAG chunk. The public free endpoint returns one page's Markdown in JSON. The application chooses boundaries, stores chunk records, builds the index, and decides when a source needs refreshing.

Try a heading-aware split on one public page

The free URL-to-Markdown tool exposes the one-page JSON route used below. The example requests IANA's public reserved-domains page, checks the returned page status, discards navigation before the first heading, and keeps the current H1/H2 path with each section. It uses Python's standard library and needs no AnyCrawler API key.

import json
import re
import urllib.parse
import urllib.request

target = "https://www.iana.org/domains/reserved"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urllib.parse.urlencode({"url": target})
request = urllib.request.Request(endpoint, headers={"Accept": "application/json", "User-Agent": "Mozilla/5.0"})
with urllib.request.urlopen(request, timeout=30) as response:
    data = json.load(response)

if not data.get("ok") or data.get("status_code") != 200:
    raise RuntimeError("The source page was not retrieved successfully")
markdown = data.get("results", {}).get("markdown", "")
title = data.get("results", {}).get("title")
lines = markdown.splitlines()
first_heading = next((i for i, line in enumerate(lines) if re.match(r"^#{1,2} ", line)), None)
if first_heading is None:
    raise RuntimeError("No trustworthy section boundary found")

chunks = []
heading_path = []
body = []

def flush():
    text = "\n".join(body).strip()
    if text:
        chunks.append({
            "source_url": data["requested_url"],
            "title": title,
            "heading_path": heading_path.copy(),
            "ordinal": len(chunks),
            "text": text,
        })

for line in lines[first_heading:]:
    match = re.match(r"^(#{1,2}) (.+)$", line)
    if match:
        flush()
        body = []
        level = len(match.group(1))
        heading_path = heading_path[: level - 1] + [match.group(2)]
    else:
        body.append(line)
flush()

print(json.dumps({"source": data["requested_url"], "status": data["status_code"],
                  "title": title, "chunks": len(chunks),
                  "headings": [c["heading_path"] for c in chunks]}, indent=2))

In the recorded public run, the endpoint and target page both returned 200. The response title was IANA-managed Reserved Domains; the script produced six nonempty heading sections. The response did not expose final_url, canonical_url, or credits_used, so the example makes no claim about those fields or about a fresh uncached origin fetch. The free tool may serve a cached result. The split is deliberately small: it demonstrates provenance, not production-ready Markdown parsing or a measured retrieval improvement.

When headings do not make safe boundaries

A heading can introduce a table whose column labels and notes must stay together, or a code sample whose explanation is in the paragraph above it. If a section is too large, split at a semantic boundary inside it and repeat the parent heading path in each child record. When no trustworthy boundary is available, stop or mark the page for a parser designed for that document shape; a fixed character cut should not silently turn an incomplete fragment into a citable claim.

At retrieval time, show the model the chunk ID, title, heading path, source URL, and passage text together. Microsoft's RAG prompt guidance recommends labeling retrieved chunks and including document title, section heading, or source URL so the answer can point to the supporting material. The application still has to verify that each generated claim is supported by the cited passage and that the passage belongs to the source page shown to the reader.

Check a citation before accepting an answer

For each answer sentence that cites a chunk, reopen the stored snapshot or live page and locate the exact passage. Check its heading context, any nearby qualification, and whether the source version is still appropriate for the user's question. If the answer combines two sections, it may need two citations. If a retrieved chunk lacks the required context, retrieve a neighboring section or abstain on that part; adding a plausible URL after generation does not repair an unsupported sentence.

The open question is how long a chunk should remain citable after its source changes. A product FAQ may tolerate a slower refresh than a status page, but no universal interval follows from the page format. Choose a refresh and retention rule for the reader's task, keep versioned snapshots where appropriate, and evaluate whether old chunks are removed from the index when the source changes.

Frequently asked questions

Should I always chunk a web page by headings?

Headings are a useful first boundary when the extracted Markdown preserves the page's real structure. They do not guarantee that each section is small enough for an embedding model or complete enough to support a citation. Inspect tables, code blocks, notes, and long sections before indexing. If a heading is missing or misleading, use a parser that understands the page shape or hold that page for review rather than pretending a character cut preserves meaning.

What source metadata should every RAG chunk keep?

At minimum, keep the page URL you can actually open, the page title, the chunk's heading path, its position, and the exact text that was indexed. Retain retrieval time and a snapshot identifier when source versions matter. If the fetch response supplies requested, final, and canonical URLs, store them as separate fields rather than replacing one with another. That record lets a reviewer locate the passage and understand which page state supported an answer.

Can I cite the URL returned with a retrieved chunk without checking it?

The URL is a route to evidence, not proof that the answer sentence follows from the passage. A chunk can be stale, incomplete, taken from the wrong section, or stripped of a qualification. Check the quoted or paraphrased claim against the stored page snapshot and its heading context. For time-sensitive material, compare the snapshot with the current page and decide whether to refresh the index before giving the answer.

Does AnyCrawler split and index the page for my RAG system?

The public free tool returns one accessible page as Markdown inside a JSON response. It does not expose a complete RAG ingestion or citation-validation workflow. Your application chooses how to split the Markdown, stores chunk metadata, creates embeddings or another search index, retrieves passages, and checks citations. The example here uses the free route to demonstrate a reproducible one-page input; production integrations can use the documented page extraction API when they need explicit Fetch or Render control.

What if the extracted page has no usable headings?

Treat missing headings as a quality signal, not an invitation to invent a source structure. First verify that the expected body was fetched; an empty JavaScript shell or navigation-only result may require a different access path or a manual check. If the body is complete but unstructured, choose boundaries suited to that page, such as paragraphs or records, and keep the source and position metadata. Do not present an arbitrary split as an exact section citation.