To extract tables from web pages reliably, keep each value attached to its row label, full column-header path, units, notes, and source before converting it into a spreadsheet or JSON record. Select the intended table explicitly, preserve the original cell text, and check ambiguous relationships against the source. A successful page request only establishes that content arrived; it does not prove that merged headers or missing values survived conversion. Start with a small, known table, reject unexpected schema changes, and retain enough evidence to explain every normalized value later.

Decide what a cell must mean outside the page

Imagine a table whose cell says 12. That value becomes useful only after you know whether it describes shipments, dollars, a percentage, or a count, and which entity and period it belongs to. A nearby note might restrict it to one region or exclude a product category. Once the cell leaves the page, those relationships need an explicit place to live.

Define that context before choosing a parser. The following is a proposed extraction contract, not an industry score or a claim about a particular API.

Preserve Why it changes interpretation Acceptance check
Table identity and caption A page can contain several unrelated tables The chosen table matches the requested subject
Row label and column-header path Repeated labels can belong to different groups Each value has an unambiguous row and full header ancestry
Original cell text Formatting, symbols, and identifiers can be meaningful Raw text remains available beside any typed value
Units, period, and qualifying notes The same number can describe different quantities Required context is present or explicitly unknown
Source URL and saved observation The page may change after extraction A reviewer can locate the table in the retained source

This contract can be smaller for a simple directory and stricter for a comparison involving quantities. Choose it from the downstream question. Collecting every possible field is unnecessary, but silently dropping a field that changes the answer makes the record unreliable.

Acquire the page, then inspect what actually arrived

When the URL is known, read the page and confirm that its useful table content is present. Fetch is appropriate for content already in the returned HTML. Render can help when JavaScript must run before the table appears. Search helps find candidate URLs; Screenshot can preserve visible arrangement for review. Table selection, parsing, validation, and storage remain application responsibilities. Rendering a page does not imply clicking through pagination or visiting linked pages.

Use AnyCrawler's free URL-to-Markdown tool to inspect the public response before building an adapter. The following request was run against MDN's HTML table reference:

curl --fail --get \
  --data-urlencode 'url=https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/table' \
  'https://api.anycrawler.com/free/v1/crawl' \
  --output table-page.json

The checked response returned gateway HTTP 200, ok: true, target status_code: 200, and Markdown containing the reference page. It did not return final_url or credits_used; preserve those as unknown rather than inventing values. The Free route may return cached results, so distinguish when your application received the response from the response's own timestamp.

That Markdown also contained navigation and example code. Finding a header word in it would not prove that a specific displayed table had been recovered. Inspect the relevant region, distinguish a code sample from the table it describes, and compare its cell relationships with the original. This probe tests public acquisition; it does not establish the authenticated Fetch or Render response contract or universal table-extraction quality.

Keep grouped headers until you choose a schema

Flattening repeated headers too early can erase the distinction between groups. W3C's irregular-header tutorial shows a table where the same lower-level measures belong to different upper-level headers. A useful representation keeps a path such as group → measure, rather than renaming duplicate columns with unexplained suffixes. Row-spanning labels deserve the same attention.

When you have separately obtained permitted HTML, a table parser can retain more structure than a plain-text split. The pandas read_html documentation describes table selection, multiple header rows, and handling of spanning cells, while warning that cleanup may still be necessary. It does not promise that an arbitrary page's semantics can be inferred.

This Python example was run with a saved copy of the W3C tutorial HTML, named w3c-irregular.html, using pandas and Beautiful Soup. The source file was downloaded separately; AnyCrawler's Free response is Markdown, not this raw HTML input.

from io import StringIO
from pathlib import Path
import pandas as pd
from bs4 import BeautifulSoup

html = Path("w3c-irregular.html").read_text(encoding="utf-8")
soup = BeautifulSoup(html, "html.parser")
tables = [t for t in soup.find_all("table")
          if {"Mars", "Venus"} <=
          {h.get_text(strip=True) for h in t.find_all("th")}]
if len(tables) != 1:
    raise ValueError("Expected one known tutorial table; review the source")

frame = pd.read_html(StringIO(str(tables[0])), header=[0, 1])[0]
expected = [("Mars", "Produced"), ("Mars", "Sold"),
            ("Venus", "Produced"), ("Venus", "Sold")]
if list(frame.columns[1:]) != expected:
    raise ValueError("Header hierarchy changed; do not guess a mapping")

for _, row in frame.iterrows():
    print({"row_label": str(row.iloc[0]),
           "values": [{"header_path": list(column),
                       "value": int(row[column])}
                      for column in expected]})

The check retained the expected grouped columns. The unlabeled first column became parser-generated Unnamed labels, so the example assigns it a row_label role only after inspecting this known source. Its integer conversion is suitable for these tutorial counts, not a general rule for currencies, identifiers, percentages, or blank cells. Keep the original HTML alongside the parsed result. These assertions make a changed table fail visibly; they do not make this a universal parser.

Retain raw values before normalizing them

Store the original string before parsing numbers. A leading zero may belong to an identifier; a dash may mean unavailable; a percent sign changes the interpretation of a magnitude. Decide those meanings from the table's notes and your schema. A blank cell should not silently become zero, and a missing unit should remain unknown until the source establishes it.

A practical record can contain table_id, row_label, header_path, raw_text, normalized_value, unit, period, note_ids, source_url, observed_at, and snapshot_hash. This is a suggested application schema, not an AnyCrawler response format. Include only the fields your task needs, but keep raw evidence and derived values distinguishable. If a qualifier lives outside the table, retain its text or an exact pointer into the saved page.

For retrieval, attach this context to the extracted row or cell so it remains understandable when separated from neighboring rows. The page URL identifies where the observation came from; a snapshot hash identifies the bytes you saved. Neither tells you that the source was correct, and a response timestamp alone does not establish when its underlying content last changed.

Make failures change the next action

An access failure should stop the extraction path. During the same checks, the Free request for the W3C tutorial returned HTTP 502 with an error saying its robots.txt retrieval received 403. That attempt did not produce table data, and changing to Render would not establish permission to bypass the failure. The successful MDN request above is a separate observation, not a repaired W3C extraction.

When access succeeds but the intended table is absent, check whether you selected the wrong document region, received an incomplete page, or need JavaScript execution. When the cells arrive but grouped headers have disappeared, obtain an appropriate permitted source representation or send the record to review. Guessing a group from cell position conceals the uncertainty.

For paginated or filtered tables, record the visible page and active scope. A partial table can be useful if labeled honestly; it cannot support a claim about the entire dataset. If a column or note changes, reject the old schema until the new relationship is understood. Keep failed and review-required records separate from accepted records so downstream calculations cannot quietly include them.

When does a table need its own adapter?

A recurring decision may justify a parser maintained for one table family, with explicit header paths and change checks. Occasional or highly irregular tables may be better handled through review with a saved visual and textual source. A renamed measure or a revised footnote can change a table's meaning even when its layout stays familiar. Before expanding automation, ask which schema changes the application can detect and who will resolve a result that remains ambiguous.

Frequently asked questions

Is Markdown enough to preserve a web table?

Markdown can be useful when the table has straightforward rows and columns and the conversion retains the needed labels. Inspect the specific output before relying on it. Grouped headers, row spans, notes outside the table, and repeated labels may need a richer representation or source review. Keep the original observation and validate the relationships required by the downstream question. A document that contains the right words can still assign them to the wrong cells.

How should duplicate column names be handled?

Check whether the names belong to different parent headers before renaming them. Preserve the full header path, such as a group followed by a measure, in the normalized record. If the source has no clear relationship that distinguishes the columns, mark the mapping for review. Adding numeric suffixes can make a DataFrame technically usable while leaving its meaning unresolved. A known table adapter should reject unexpected header paths rather than silently accept a new arrangement.

Should an empty table cell become zero?

Only when the source and the application schema explicitly define that meaning. Blank, unavailable, suppressed, and zero can represent different states. Preserve the original cell text and record a separate normalized value or missing-value status. The same caution applies to dashes, symbols, percentages, and identifiers with leading zeros. If a note explains the convention, carry that note with the record so a later reader can understand why the conversion was accepted.

Does browser rendering solve every table-extraction problem?

Rendering can make content available when JavaScript is required before it appears. It does not by itself resolve ambiguous headers, identify a unit, validate a number, or prove that pagination is complete. Start by determining whether the missing content is a retrieval problem or a parsing problem. Keep the visible scope explicit and respect access failures. Table interpretation and acceptance checks remain application work after the page content has been obtained.

What should an agent cite for a value extracted from a table?

Keep the source page URL together with the table identity, row label, full column-header path, and relevant notes. Retain a saved observation and its collection time so a reviewer can locate the value even if the live page changes. A screenshot may help explain visual grouping, while the textual record supports processing. Cite only the relationship you verified, and leave unknown provenance fields explicit instead of substituting a guessed final URL or timestamp.