To extract JSON-LD from a web page, read its HTML, collect the application/ld+json script blocks, and parse each block while preserving its source and any errors. Then identify the entity your task concerns and compare its fields with the page's visible content. A successful JSON parse tells you what the publisher encoded; it does not verify the claim. Keep the raw markup and readable text as separate evidence, especially when dates, prices, or authors disagree. The workflow below inventories a permitted public page and shows where an agent should pause for review.
What the script tells you—and what it leaves unproved
JSON-LD gives a page a machine-readable description. A blog page might describe the article, its author, the publishing organization, and the navigation trail. Those records can all be valid while answering different questions. Reading the first script and assuming it represents the main article can therefore select the wrong object.
The W3C JSON-LD specification's HTML embedding rules describe these scripts as data blocks. For an initial inventory, retain each block separately so you can trace a field back to the markup that supplied it. Full JSON-LD processing has additional rules for contexts, identifiers, and combining embedded data; the simple inventory in this article does not implement that semantic processing.
Even the right object contains publisher assertions. An article date might reflect a template default. A product offer might describe another variant. A review count might disagree with what readers can see. Treat the extracted value as a candidate field until the surrounding page supports the meaning your application assigns to it.
Keep HTML and readable content as separate inputs
Use a public page you are permitted to read, check its access rules, and preserve the response before converting or normalizing it. The HTML capture supplies script blocks. The readable body supplies the context used to corroborate them. Save the requested URL, final URL, response status, capture time, and a hash of the original bytes with each capture.
AnyCrawler's free URL-to-Markdown tool lets you inspect the readable-text side without an API key. Its output is a JSON response containing Markdown. That output contract does not promise the original HTML or a JSON-LD inventory, so obtain raw HTML separately when your task requires embedded scripts. For a reproducible public example, this request reads AnyCrawler's own Markdown guide:
curl --get 'https://api.anycrawler.com/free/v1/crawl' \
--data-urlencode 'url=https://anycrawler.com/blog/webpage-to-markdown/' \
--output readable-page.json
The check of this page returned HTTP 200, a target status_code of 200, and a title, description, and readable results.markdown. That free response did not return final_url, canonical_url, or credits_used. Record those fields as unknown when absent; a final URL observed in your separate HTML request belongs to that request, not automatically to the API response.
The free tool may serve a cached result. Its response timestamp and your request time are useful records, but two independent requests do not create a synchronized snapshot. If a comparison depends on the current price or another changing value, check whether both captures describe the same page state before judging a difference.
Inventory every block with Python
This standard-library example reads the same public page, limits the response size, and records successful parses alongside malformed blocks. Save it as inventory_jsonld.py and run python inventory_jsonld.py. The timeout and byte cap are sample safeguards chosen for this example; adapt them to your permitted inputs. Network errors stop the script so they cannot be mistaken for a page with no structured data.
import hashlib
import json
from datetime import datetime, timezone
from html.parser import HTMLParser
from urllib.request import Request, urlopen
class JsonLdBlocks(HTMLParser):
def __init__(self):
super().__init__()
self.blocks = []
self.current = None
def handle_starttag(self, tag, attrs):
kind = (dict(attrs).get("type") or "").split(";")[0]
if tag == "script" and kind.strip().lower() == "application/ld+json":
self.current = []
def handle_data(self, data):
if self.current is not None:
self.current.append(data)
def handle_endtag(self, tag):
if tag == "script" and self.current is not None:
self.blocks.append("".join(self.current))
self.current = None
url = "https://anycrawler.com/blog/webpage-to-markdown/"
request = Request(url, headers={"User-Agent": "AnyCrawlerEditorialCheck/1.0"})
with urlopen(request, timeout=30) as response:
if response.headers.get_content_type() != "text/html":
raise ValueError("Expected an HTML response")
raw = response.read(2_000_001)
if len(raw) > 2_000_000:
raise ValueError("HTML response exceeds this example's byte cap")
source = {"requested_url": url, "final_url": response.url,
"http_status": response.status,
"captured_at": datetime.now(timezone.utc).isoformat(),
"html_sha256": hashlib.sha256(raw).hexdigest()}
html = raw.decode(response.headers.get_content_charset() or "utf-8")
parser = JsonLdBlocks()
parser.feed(html)
parser.close()
records = []
for index, block in enumerate(parser.blocks):
record = {"block_index": index, "raw_jsonld": block}
try:
record.update(status="parsed", value=json.loads(block))
except json.JSONDecodeError as error:
record.update(status="invalid_json", error=str(error))
records.append(record)
print(json.dumps({"source": source, "blocks": records}, indent=2))
The implementation uses Python's HTMLParser handlers to collect script content without executing it. In the controlled page check, the HTML request returned 200 and the inventory found two parseable blocks: BlogPosting and BreadcrumbList. That establishes the example's behavior on this page. It does not establish extraction coverage across arbitrary websites.
The script keeps the full parsed JSON value. It does not implement the JSON-LD processing algorithms for expanding @context, following remote contexts, merging graphs, or resolving keyword aliases, and it does not validate a Schema.org type. A top-level @graph can contain several entities. Inspect that structure before selecting fields; do not flatten it by taking whichever name or datePublished appears first. If an opening script never closes, this collector does not finalize it; preserve the HTML and classify that input for inspection.
The printed hash identifies the captured HTML bytes. Store the HTML itself if you need to reproduce the extraction later. A hash without the underlying capture cannot show an editor what the page contained, and neither the hash nor a successful parse proves that the publisher's statement was true.
Select the entity before accepting a field
Choose an object based on the reader task, not just a convenient field name. For an article ingestion job, check the article object against the page's headline, URL identity, byline, and date context. For an offer, identify the relevant product or variant before interpreting currency, price, or availability. Keep the raw value when normalization would hide an ambiguity.
This acceptance table separates common decisions an ingestion pipeline needs to make:
| Observation | Record it as | Next action |
|---|---|---|
| Correct entity, field supported by visible context | Corroborated publisher field | Save the value and the supporting passage or capture location |
| Parseable object, entity identity unclear | Unresolved entity | Inspect identifiers, links, nesting, and the page's main subject |
| Expected field absent | Missing field | Leave it unknown; use an explicit fallback only with its own evidence |
| Several objects provide conflicting values | Conflicting candidates | Retain each value and block index; route the conflict for review |
| Markup and visible text disagree | Uncorroborated field | Check capture timing, variant, region, and whether the template is stale |
| JSON parsing fails | Invalid block | Save the raw block and parser error; continue inventorying other blocks |
Google's general structured-data guidelines require markup to represent the visible page content and warn that technically valid markup does not guarantee a search result feature. That guidance supplies a useful corroboration check. It does not turn an extractor into a validator of publisher honesty or a predictor of search visibility.
For each accepted application field, preserve the selected block index, the object's identifier when present, the original value, the normalized value, and the reason for accepting it. Attach the readable capture used for comparison. If the evidence supports only an article headline, do not label the whole object “verified”; status belongs to the specific field and check you performed.
Handle missing and conflicting markup
A zero-block inventory means no qualifying closed script was collected from that HTML response. It does not prove that the page has no structured data. The page could use another markup format, or JavaScript could insert JSON-LD later. Google's JavaScript structured-data guide documents dynamically generated markup. When that distinction matters, inspect a separately captured rendered DOM through an appropriate, permitted browser workflow.
AnyCrawler's Fetch and Render text outputs, a raw HTML response, and a DOM snapshot are different artifacts. Search helps discover URLs; Screenshot can preserve visible state; neither substitutes for script inventory. Login, clicking, submitting forms, and other interactions require a suitable automation workflow. This article's example performs a single public HTML request and a separate free Markdown request, with no recursive crawl or scheduled monitoring.
Keep malformed blocks visible in your result instead of quietly repairing them into plausible JSON. A trailing comma, truncated script, or unexpected scalar should remain a parse or shape issue. If a page has both old and new offer objects, do not silently resolve the conflict by script order. Compare what each object identifies and whether the visible page explains the difference.
Time can also explain a disagreement. A cached text extraction and a fresh HTML response may represent different revisions. Preserve both times and acquisition methods, then obtain comparable evidence if the decision requires it. An empty value or stale candidate is safer to review than a confident field assembled from mismatched captures.
How much corroboration does your task need?
A catalog index may use a publisher-supplied headline after a straightforward identity check. A claim about a changing offer needs more context: the applicable variant, conditions, and captured page state. A claim whose truth matters beyond describing the publisher may require independent evidence. Decide that standard before the extracted field enters your agent's answer or triggers an action.
The open design question is how much ambiguity your pipeline may carry forward. Rejecting every imperfect block can discard useful sources; accepting every parseable value can disguise contradictions. Choose which fields can remain unknown, which conflicts require human review, and which claims need another source. Then make those decisions visible in the evidence record, so a downstream model cannot turn an unresolved candidate into a settled fact.
Frequently asked questions
Can I extract JSON-LD from a Markdown response?
Only if that response explicitly preserves the original JSON-LD content. Readable Markdown commonly serves a different purpose: presenting the page body for review or model context. AnyCrawler's free tool returns Markdown inside JSON, which does not mean the JSON response contains embedded structured-data scripts. Keep a separate HTML capture for the script inventory and use the readable response to check the surrounding claims. Record the acquisition time and method for both.
Why does a page contain more than one JSON-LD script?
A page can describe several things, such as an article, a publishing organization, and a breadcrumb trail. Separate blocks can therefore be useful rather than duplicate noise. Inventory the blocks before choosing the object relevant to your task, and retain the selected block's provenance. A full JSON-LD processor has semantic rules for handling embedded data together; the example here preserves blocks independently so an editor can inspect what each one supplied.
Is json.loads enough to validate JSON-LD?
It checks whether the block is valid JSON syntax. It does not interpret remote contexts, expand terms, resolve keyword aliases, merge graphs, or confirm that an object follows the vocabulary rules your application expects. Use a suitable semantic processor when those operations matter. Even semantic validity would not prove the publisher's claims. Keep separate statuses for parsing, entity selection, field shape, and corroboration against the page or other evidence.
What should I do if the HTML contains no JSON-LD?
Keep that result limited to the captured response. Confirm that the request succeeded and returned the intended page before interpreting an empty inventory. The page may use another structured-data format or add JSON-LD through JavaScript. If your task requires checking that possibility, inspect a permitted rendered DOM separately. Do not assume a text extraction API exposes that DOM, and do not manufacture missing fields from unrelated objects or page snippets.
Which value wins when JSON-LD and visible text disagree?
Neither should win automatically. Keep both candidates, their source locations, and their capture times. Check whether they describe the same entity, variant, region, and page revision. A cached readable response can differ from a fresh HTML capture, while a template can contain an outdated field. If those checks do not explain the difference, mark the field unresolved and request review or additional evidence before using it in a consequential decision.






