Choose Markdown when an agent must read, summarize, quote, or chunk a page; choose structured JSON when downstream code requires named fields and deterministic validation; keep HTML when DOM structure, attributes, embedded metadata, or a later reprocessing step matters. These formats solve different problems, so the strongest pipeline often stores a faithful source artifact and derives a task-specific representation from it. Decide from the operation that follows extraction, then verify that the chosen form preserved every field, relationship, and provenance signal that operation needs.

Start with the operation, not the file extension

The useful question is not which format is generally “best for AI.” It is what the next component must do without guessing.

If a model needs to understand an article, Markdown removes much of the navigation and presentation noise while keeping headings, paragraphs, lists, links, tables, and code in a compact reading order. If application code must route a record, compare a price, or reject a missing date, structured JSON gives those values stable names and types. If a later stage may need a data attribute, canonical link, table cell relationship, or embedded machine-readable block, the original HTML remains the safer source artifact.

That makes format choice a contract decision:

Downstream task Preferred working format What must remain available Primary failure signal
Summarization, quotation, or semantic chunking Markdown headings, paragraph order, links, code, tables, source URL required section or link disappears
Field extraction, routing, comparison, or alerts Structured JSON schema, types, null policy, provenance for each field missing or invented field passes validation
DOM analysis, metadata recovery, or future reprocessing HTML elements, attributes, embedded data, original source bytes conversion removes information needed later
Auditable research or high-value ingestion HTML plus a derived form immutable source, transformation version, output, validation record derived output cannot be traced to its source

The three choices are not equal alternatives at every stage. HTML is usually the captured document, Markdown is a readable transformation, and structured JSON is an application-owned interpretation. Treating all three as interchangeable hides where information was removed or judgment was introduced.

HTML is the fidelity layer

HTML carries relationships that plain text cannot fully express: element hierarchy, link destinations, table structure, language annotations, data-* attributes, canonical links, and embedded JSON-LD. The WHATWG HTML Living Standard defines both document semantics and the parsing model browsers use to create a DOM. Keeping HTML therefore preserves options for later selectors, re-extraction, and audits.

That fidelity has a cost. A production page may contain navigation, consent controls, scripts, style data, repeated templates, hidden content, and markup that has little value to the current model task. Sending it all to a model can blur the document boundary and makes it harder to prove which visible or machine-readable part supported an answer.

Keep HTML when the transformation is not yet trustworthy, when the task depends on DOM relationships, or when you expect to improve the extractor later. Store it with the requested URL, final URL when known, retrieval status, content type, capture time, and a byte hash. Those fields establish identity; the HTML alone does not.

Markdown is the reading layer

Markdown is strongest when document structure should remain legible to both people and language models. The CommonMark specification defines a plain-text structured-document syntax for blocks and inline content. In a web extraction pipeline, the conversion can retain useful headings, paragraphs, lists, links, tables, and fenced code while dropping scripts and much presentation markup.

Conversion is still a lossy operation. A Markdown heading does not preserve every HTML attribute. A flattened table may lose header associations or nested layout. Content selection can omit a disclosure, sidebar, or repeated label that the task actually needed. Before adopting Markdown, run a web-page-to-Markdown conversion and QA workflow against pages that resemble your production inputs.

Use an explicit Markdown contract instead of checking only that the string is non-empty. For an article-ingestion job, that contract might require the title, two known section headings, at least one source link, all fenced code blocks, and a minimum set of phrases. A short but correct page should pass; a long navigation dump should fail.

Structured JSON is the control layer

JSON gives software named fields, arrays, primitive values, and nested objects. RFC 8259 defines JSON as a text format for serialized structured data, including objects, arrays, strings, numbers, booleans, and null. Those mechanics are useful only after the application defines what each field means.

For example, a price record should state whether amount is required, which currency syntax is accepted, whether a billing period can be null, and where the value came from. Without that schema, valid JSON can still contain the wrong price, a guessed currency, or a title copied from an unrelated card.

Structured output should therefore fail closed. If a required field is absent, return a validation error or explicit null according to the contract. Do not ask a model to fill the gap from background knowledge and then present the result as extracted page data. Preserve source snippets, selectors, or offsets beside high-value fields when reviewers need to trace them.

One page in three representations

The following example was run against the permitted public page https://example.com/ on September 6, 2026. The source request and the AnyCrawler request both returned HTTP 200, and the extracted Markdown contained the expected documentation-domain passage. The test then created a small JSON projection from fields that had passed explicit checks. It did not test model accuracy, JavaScript rendering, latency, cost, or other websites.

const sourceUrl = "https://example.com/";
const crawlUrl = new URL("https://api.anycrawler.com/free/v1/crawl");
crawlUrl.searchParams.set("url", sourceUrl);

const [sourceResponse, crawlResponse] = await Promise.all([
  fetch(sourceUrl),
  fetch(crawlUrl),
]);

if (!sourceResponse.ok || !crawlResponse.ok) {
  throw new Error("A required read failed");
}

const sourceHtml = await sourceResponse.text();
const crawl = await crawlResponse.json();
const markdown = crawl.results?.markdown;
const requiredText = "This domain is for use in documentation examples";

if (crawl.ok !== true || crawl.status_code !== 200) {
  throw new Error("The page result failed");
}
if (typeof markdown !== "string" || !markdown.includes(requiredText)) {
  throw new Error("The extracted document is incomplete");
}
if (!sourceHtml.includes(requiredText)) {
  throw new Error("The source artifact does not match the fixture");
}

const record = {
  title: crawl.results?.title ?? null,
  source_url: crawl.final_url ?? crawl.requested_url ?? sourceUrl,
  status_code: crawl.status_code,
  required_passage_present: true,
};

console.log(record);

The HTML is the replayable source. The Markdown is the reading representation. The JSON object is the narrow contract used by downstream code. It would be misleading to call that projection a complete representation of the page: it intentionally drops the body, presentation, links, and most metadata.

AnyCrawler's current Fetch Web Page API contract follows the same separation. A known public URL can return normalized Markdown together with page identity, status, metadata, links, and usage fields when requested or available. Your application still owns the schema for business-specific fields and the validation that turns page content into those fields.

Validate the representation you actually consume

A successful HTTP response proves transport, not content quality. Each representation needs checks tied to the downstream task.

For HTML, verify the final URL, content type, status, expected DOM landmarks, canonical relationship when used, and a hash of the captured bytes. If useful content is absent from the initial HTML, that is a render decision rather than permission to accept an empty shell.

For Markdown, check the document boundary and the structures the model will use. Required headings, links, code fences, table rows, or known phrases are stronger signals than length alone. Record the extractor and configuration because a different readability rule can change the output without the page changing.

For JSON, validate against a schema and distinguish missing, null, invalid, and unsupported. Reject unknown enum values when they would change routing. Keep provenance beside claims that may be reviewed. A syntactically valid object is not evidence that its fields came from the correct page region.

A compact validation record can make those decisions auditable:

{
  "source": {
    "requested_url": "https://example.com/",
    "status_code": 200,
    "html_sha256": "..."
  },
  "transform": {
    "format": "markdown",
    "version": "readability-v3",
    "output_sha256": "..."
  },
  "checks": {
    "required_passage_present": true,
    "schema_valid": true
  }
}

Combine formats without creating silent drift

Storing more than one representation raises a new problem: derived artifacts can disagree. A crawler may refresh HTML while a stale Markdown cache remains. A schema change may rename a JSON field without regenerating older records. A reviewer may see a value that no longer matches the preserved page.

Connect every derived artifact to the source hash and transformation version. Regenerate Markdown or JSON when either changes. If freshness matters, record when the page was read separately from when the transformation ran. When an application serves cached output, make that state visible to the component deciding whether the evidence is fresh enough.

For high-value jobs, keep a simple lineage:

  1. Capture the requested URL, resolved identity, status, timestamp, and source bytes.
  2. Hash the source artifact.
  3. Transform it using a named extractor and configuration.
  4. Validate task-specific content requirements.
  5. Store the derived hash and any field-level provenance.
  6. Send only the representation the next component needs, while retaining a path back to the source.

This pattern limits model context without deleting the evidence needed for repair or review.

Diagnose format failures by where information was lost

When an agent answer is wrong, changing the prompt is not always the right first move. Locate the earliest broken contract.

  • If the initial HTML lacks the required content, check whether the page needs JavaScript rendering or whether access failed.
  • If HTML contains the content but Markdown does not, inspect content selection and structure conversion.
  • If Markdown contains the value but JSON does not, inspect the extraction prompt, parser, schema, and validation rules.
  • If all stored artifacts are correct but the answer is wrong, then investigate retrieval, chunk selection, instructions, and model behavior.

This order avoids compensating for an extraction defect with a larger context window. It also tells you which artifact to reproduce when a failure appears only on certain page types.

Should one representation be the canonical record?

There is no universal storage answer because fidelity, privacy, retention cost, and reprocessing needs vary by application. Keeping raw HTML supports future extraction, but it may preserve irrelevant or sensitive material that a narrow task does not need. Keeping only Markdown reduces noise, but a later parser cannot recover attributes or omitted regions. Keeping only JSON minimizes ambiguity for downstream code, but locks the record to today's schema and extraction judgment.

The open design question is how much reversibility the workflow needs. Define retention and access rules for the source separately from the small representation sent to a model. Then test whether you can explain a derived field, regenerate it after an extractor change, and delete source material on schedule. Those operational checks matter more than declaring one syntax the permanent winner.

Frequently asked questions

Is Markdown always better than HTML for LLM context?

No. Markdown is often easier to read and chunk because it removes much presentation markup while retaining document structure. HTML remains useful when the task depends on DOM hierarchy, attributes, canonical links, embedded metadata, or a later extraction pass. Choose Markdown for a verified reading contract, and retain HTML when losing source structure would make the result hard to audit or regenerate.

When should an agent receive structured JSON instead of Markdown?

Use structured JSON when the next step expects named fields, stable types, explicit null behavior, or deterministic routing. Examples include comparing product records, rejecting a missing publication date, or triggering a workflow from an approved enum. Define and validate the schema first. JSON syntax alone does not prove that values are complete, correct, or grounded in the intended page region.

Should I store both the source HTML and extracted Markdown?

Store both when the content is valuable enough to reprocess, audit, or debug later. Link the Markdown to a hash of the exact HTML and record the extractor version and configuration. For low-risk, disposable tasks, that retention may be unnecessary. Make the choice from reversibility, privacy, storage policy, and the cost of fetching the source again, rather than from model preference alone.

Does fewer context tokens mean a format will produce a better answer?

No. A smaller representation can reduce processing cost and irrelevant material, but it may also remove the table, disclosure, link, or relationship needed for the task. Token counts also depend on the tokenizer and exact input. Evaluate output with task-specific completeness checks and answer quality on representative pages. Do not use character reduction or token reduction as a substitute for those tests.

Can AnyCrawler return arbitrary structured JSON fields from any page?

The public page-reading contract returns normalized Markdown plus page identity, status, metadata, links, media, and usage fields when requested or available. A business-specific object such as a price record or policy record is an application-owned projection. Define its schema, validate required fields, preserve provenance, and return a clear failure when the page cannot support the requested structure instead of inventing values.

What should I record so a transformed page remains auditable?

Record the requested URL, final URL when known, retrieval status, content type, capture time, and a hash of the source artifact. For each derived form, add the transformation name and version, configuration, output hash, and task-specific validation results. High-value structured fields may also need source snippets, selectors, or offsets. This lineage lets a reviewer reproduce where information changed or disappeared.