Evaluate web extraction quality against the task the LLM must perform: check that required facts and sections survive, irrelevant page furniture stays out, tables and code retain their meaning, and the document carries usable source information. Make those checks explicit before admitting content to a prompt or retrieval index. A successful HTTP request and a shorter document are useful observations, but neither establishes completeness. Use a reference page and deliberate failure examples to test the checks, preserve the original extraction, and withhold any result that cannot support the intended answer.
Define what the model needs before judging the extraction
Start with a concrete question. If the model must explain an API parameter, the input needs the parameter name, allowed values, defaults, and exceptions. If it must compare plans, the input needs the plan names, prices, currency, billing period, and qualifying footnotes. A clean title and an introductory paragraph cannot substitute for those details.
Inspect a representative source page and record the material that must survive extraction. Keep a reference snapshot or reviewed excerpts so a future test compares the same document state. Write down the task, expected sections or fields, unwanted regions, and any required relationships. A reference should identify what the page actually contains; it should not be an LLM reconstruction of what you think it ought to say.
This makes completeness and noise separate questions. Trafilatura's extraction documentation describes settings that favor precision or recall: aggressive cleaning can reduce unwanted material while losing useful text, and broader retention can recover content while keeping more noise. The relevant choice depends on which omissions would change your answer. These are general extraction tradeoffs, not a claim that AnyCrawler uses that library.
Use a scorecard with reasons, not an unexplained total
The following scorecard is a proposed review contract. Each row produces a pass, fail, or review decision with evidence. There are no weights or universal passing percentages. Decide which rows are mandatory for your task before running the extraction; an optional missing author is different from a missing price qualifier.
| Dimension | What to inspect | A failure that matters | Next action |
|---|---|---|---|
| Completeness | Required sections, fields, exceptions and final paragraphs | Parameter default exists in the source but is absent from the extraction | Locate the missing region and repair retrieval or extraction |
| Boilerplate | Navigation, repeated menus, unrelated cards and consent text | Sidebar options dominate a documentation answer | Isolate the document body and recheck required content |
| Structure | Heading ancestry, list order, table rows and code indentation | A value is attached to the wrong plan or parameter | Preserve relationships in Markdown, HTML or explicit records |
| Links | Required link text and its actual destination | Citation text survives but points to an unrelated page | Compare the extracted link with the source |
| Identity and metadata | Requested/final URL, title, version and relevant dates | A redirect produces a sign-in page with a successful status | Reject the wrong document or require review |
| Task relevance | The evidence needed for the current question | A valid product overview is used to answer a missing pricing question | Find the appropriate document or report the gap |
| Freshness | Response time, known cache state and task time requirement | An old snapshot is presented as a current observation | Obtain appropriately current evidence or label its age |
Keep each failed requirement in the result. A boolean without a reason forces the next component to guess whether it should retry, render, clean, or stop. A useful result says, for example, that the parameter section is present but its default-value row is missing. It also identifies the version of the checks that produced that judgment.
Metadata needs its own review. A title can be correct while the body belongs to a navigation shell. A missing publication date can mean the source never supplied one; it does not automatically mean extraction failed. Preserve missing fields as unknown and distinguish “not present on source” from “present on source but lost.”
Check relationships as well as matching words
A phrase check can pass because a term appears in a sidebar or table of contents. For documentation, inspect the relevant section body and its code example. For a table, compare the heading-to-cell relationship rather than merely checking that every price appears somewhere. For a procedure, confirm that the steps remain in order and warnings remain attached to the relevant step.
Normalize whitespace when comparing prose, but avoid transformations that conceal damage. Lowercasing code, stripping units, removing minus signs, or sorting every line can make corrupted output look equivalent. Keep the raw extraction so you can inspect what the converter actually produced before normalization.
Build an evaluation set around your real page types. A short landing page, long documentation page, product table, discussion, and client-rendered view exercise different requirements. The Web Content Extraction Benchmark, by Murrough Foley, uses annotated reference content together with snippets that should be included and snippets that should be excluded. That inclusion/exclusion approach is useful for a local regression set. Its benchmark results do not establish a passing threshold for your own pages or your downstream model.
A short page can pass while a long response needs cleaning
In the public AnyCrawler free-endpoint probes prepared for this article, Example Domain returned a target status of 200 and 167 Markdown characters. Its purpose statement and IANA link were present. The GitHub documentation page “Organizing information with tables” also returned a target status of 200, with 20,556 Markdown characters. It retained the expected table topic but also included sidebar material such as “Start your journey.”
These are observations from individual responses, not accuracy measurements. The short page supports a narrow purpose-statement task. The longer documentation response still needs body isolation before it becomes focused model context. Neither character count gives a general quality score. The free responses did not expose final_url or credits_used; those values remain unknown.
Try a public page in the free URL-to-Markdown inspection tool and compare its output with the visible document. Successful free responses may be cached for up to 24 hours, so record the service timestamp separately from the time your program requests the result. A request made now does not prove the origin page was retrieved now.
The Python example below runs a small fixture check against the public endpoint without an API key. It then makes two deliberately damaged copies locally: one removes required text and one adds a known noise marker. The damaged copies test the validator; they are not additional crawler responses. It requires Python and curl. Save the code as check_extraction.py and run python check_extraction.py.
import copy
import hashlib
import json
import shutil
import subprocess
from datetime import datetime, timezone
from urllib.parse import urlencode
target = "https://example.com/"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urlencode({"url": target})
checked_at = datetime.now(timezone.utc).isoformat()
curl = shutil.which("curl.exe") or shutil.which("curl")
if not curl:
raise RuntimeError("Install curl before running this example")
response = subprocess.run(
[curl, "-sS", "--max-time", "60", "-w", "\n%{http_code}", endpoint],
capture_output=True, check=True,
)
body, status = response.stdout.rsplit(b"\n", 1)
gateway_status = int(status)
if gateway_status != 200:
raise RuntimeError(f"Gateway returned {gateway_status}; inspect before retrying")
data = json.loads(body)
# Fixture rules: replace these after inspecting your actual source page.
def check(record):
body = record.get("results", {}).get("markdown", "")
checks = {
"gateway_ok": gateway_status == 200,
"target_ok": record.get("status_code") == 200,
"service_ok": record.get("ok") is True,
"title": record.get("results", {}).get("title") == "Example Domain",
"required_text": "documentation examples" in body,
"source_link": "https://iana.org/domains/example" in body,
"no_test_noise": "NAVIGATION_TEST_ONLY" not in body,
}
return {
"fixture_pass": all(checks.values()),
"failed": [name for name, passed in checks.items() if not passed],
"markdown_sha256": hashlib.sha256(body.encode("utf-8")).hexdigest(),
}
missing = copy.deepcopy(data)
missing["results"]["markdown"] = missing["results"]["markdown"].replace(
"documentation examples", ""
)
noisy = copy.deepcopy(data)
noisy["results"]["markdown"] += "\nNAVIGATION_TEST_ONLY"
report = {
"target": target,
"checked_at": checked_at,
"response_timestamp": data.get("timestamp"),
"requested_url": data.get("requested_url"),
"final_url": data.get("final_url"),
"credits_used": data.get("credits_used"),
"live": check(data),
"synthetic_missing": check(missing),
"synthetic_noise": check(noisy),
}
print(json.dumps(report, indent=2))
For the response used here, the live fixture passed, removing the required phrase failed required_text, and inserting the noise marker failed no_test_noise. If your live response differs, inspect the source and returned JSON before updating the fixture. Do not loosen the checks merely to restore a green result.
fixture_pass deliberately names the limited result. The script does not validate arbitrary tables, detect every navigation block, establish freshness, or prove final-URL identity. The noise marker is a controlled negative test, not a universal boilerplate detector. The client timeout is an example bound. Example Domain is suitable for an illustrative check; use your own stable fixtures for ongoing application tests.
Route failures to the part of the pipeline that can fix them
If the requested content is absent from the initial page representation, compare it with the permitted browser view. When JavaScript supplies the missing content, follow the Fetch-to-Render diagnosis workflow. Rendering is a response to an observed retrieval gap; it does not repair every bad extraction. If the source HTML already contains the missing table, investigate selection and conversion instead.
Keep these tool boundaries explicit. Search discovers candidate URLs. Fetch reads a known page without browser rendering. Render reads a page after browser execution. Screenshot preserves a visible state for review. A task requiring login, clicks, or form submission belongs to an appropriately authorized browser automation workflow. A screenshot cannot by itself restore text or table relationships missing from model context, and a single-page endpoint does not supply a recursive crawler, scheduler, diff engine, or alerting service.
After any cleaning, chunking, or truncation, run the task checks again on the actual material sent to the model. The original extraction may contain a qualifying note that a later context limit removes. Retain the original Markdown, its hash, the admitted context or its reference, and the reasons for admission. Hashes help identify exact bytes; they do not prove those bytes are complete or true.
For recoverable network errors, use bounded retries. For a consistently wrong document or absent required content, preserve the failure and change the retrieval decision. If access remains unavailable, return an explicit evidence gap instead of asking the LLM to supply the missing material.
Who should change the contract when the question changes?
A document admitted for summarization may be inadequate for comparing a specific warranty exception. Reusing its earlier “pass” transfers an old task assumption into a new decision. Store the task and check version with the result, and re-evaluate the original evidence when either changes.
The unresolved tradeoff is how much context to retain for questions you have not anticipated. Broader retention preserves possible evidence but increases noise; narrow admission makes today's task easier to verify but may omit tomorrow's qualifying detail. Assign ownership of that decision to the application, with reviewed examples and explicit failure costs. Which missing detail would make your next answer wrong, and will the current contract notice it before the model speaks?
Frequently asked questions
Does HTTP 200 mean the extracted content is usable?
No. It shows a successful HTTP response at the layer reporting that status. A page can still contain a sign-in screen, navigation shell, or incomplete document. Check the intended title, required content, source identity, and meaningful structure separately. The example records both gateway and target checks, then tests named content requirements. Its successful fixture result only covers those requirements; production admission needs the rest of your task-specific contract.
Is there a minimum text length for good extraction?
A universal minimum would reject some useful short pages and accept some long, noisy responses. Choose expected content from the source and the reader task instead. Length can flag an unexpected change relative to a reviewed fixture, but it should trigger investigation rather than settle the decision. In this article's probes, the short Example Domain response retained its purpose statement, while the longer documentation response still contained sidebar material.
How can I test completeness without checking every page by hand?
Start with a reviewed sample covering the page types and questions your application actually handles. Record required sections, values, relationships, and excluded regions, then automate those checks and keep known failures as regression fixtures. Send uncertain or newly changed templates for review. Phrase matching can catch specific omissions cheaply, but it cannot prove complete understanding or correct relationships. Report the coverage of your checks instead of calling every passing document universally complete.
Should I render every page that fails a quality check?
Only when the failure indicates that browser execution supplies required content missing from the initial response. If the content is already present in the source but disappears during cleaning or conversion, rendering may leave the same problem. Identify whether the defect belongs to access, document selection, representation, metadata, or later truncation. Keep failed checks and the original response so the next action follows an observed cause rather than a generic retry rule.
Can an LLM judge extraction quality for me?
An LLM can assist with comparing a reference and an extraction, but its judgment should be treated as another review signal. It needs the actual source evidence and a defined task; it cannot reliably identify material that neither input contains. Keep deterministic checks for required fields and known failures, preserve reviewable evidence, and route uncertainty explicitly. A fluent assessment should not overwrite a missing-content failure or turn unknown provenance into a verified fact.






