To monitor pricing page changes without losing evidence, save each observation as a separate record containing the original response, readable text, capture time, page identity, and any screenshot. Extract the offer's currency, billing interval, unit, and eligibility before comparing it with an earlier observation. A changed number is meaningful only when those conditions match. Keep incomplete captures out of the price-change feed, and retain both versions when a reviewer accepts a change. Your application must supply the schedule, comparison logic, storage, and notifications around the page-access tools.
Define the offer you intend to compare
A pricing page can describe several different offers in the same card. A monthly equivalent under an annual commitment, a month-to-month subscription, and a temporary introductory rate should have separate identities. Saving only the large number discards the conditions that make the comparison useful.
Start with an offer key that combines the source, plan, currency, displayed unit, billed interval, region or audience, and eligibility. Keep the exact source wording beside the normalized fields. If the currency is unclear, record it as unknown instead of converting a dollar symbol into a currency code. If a plan name changes, require a reviewed mapping before treating it as the old plan with a new price.
Schema.org's unit-price vocabulary separates currency, billing duration, reference quantity, and tax information. It provides useful field names when a page supplies structured data. Still verify that the markup describes the visible offer; a machine-readable value alone does not resolve a conflicting card or footnote.
Choose a reproducible page state before collecting a baseline. Record the public regional URL, visible billing selection, relevant quantity, and whether the page was anonymous. Preserve consent banners or missing controls as capture conditions. If selecting the offer requires clicks, authentication, or form entry, use an authorized interactive browser workflow and record those actions. Rendering JavaScript does not establish that the intended controls were selected.
Keep the evidence record separate from the extracted price
Treat a capture as a stored observation and an extracted offer as an interpretation of it. That separation lets you correct a parser later without rewriting what was originally received.
| Record group | What to preserve | When to withhold comparison |
|---|---|---|
| Identity | Requested URL, returned final URL, plan identifier, region or audience | Redirect or plan mapping is unresolved |
| Time | Collector start and finish, source/API timestamp, cache metadata when available | Required freshness cannot be established |
| Offer | Exact wording, amount, currency, billing interval, unit, quantity, tax and eligibility context | A required condition is missing or differs |
| Text artifact | Original response and extracted Markdown, each with its own hash | Target card or qualification is absent |
| Visual artifact | Stored image, capture time, URL, framing and visible selected state | Image belongs to a different or unknown page state |
| Review | Extraction version, status, reviewer decision, baseline reference | Candidate change has not passed the acceptance rule |
Use a new record identifier for every attempt. Store failed attempts too, with their failure reason, so a missing observation cannot masquerade as an unchanged price. For successful reads, retain the full response before transforming text or removing navigation. A comparison copy may omit known irrelevant sections, but it should point back to the original artifact and record which transformation produced it.
Hashes help detect changes to saved bytes. They do not establish that a page was authentic, that a screenshot showed every condition, or that the offer was available to every visitor. Keep access controls and retention rules appropriate to the use of the records; choose retention deliberately before a long-running job accumulates material.
Capture the text and visible state deliberately
For a known URL whose pricing content is present in HTML, start with the Fetch page-read contract. Check for the intended plan and its qualifications before accepting the result. Use Render when the required content appears only after JavaScript executes, then apply the same completeness check. Search is useful when the source URL still needs discovery; it is not a substitute for reading the offer.
Add a visual artifact when position, a selected billing control, crossed-out text, or a nearby qualification affects interpretation. The Screenshot API's capture fields describe the separate image operation. Save the returned artifact with its own URL and timestamp, then inspect whether the relevant card and qualification are actually visible. Text retrieval and screenshot capture can observe different moments. If they disagree, mark the bundle for review rather than silently combining them into a single supposedly simultaneous observation.
Keep cache age distinct from request time. A collector can finish now while receiving a previously captured page. Where your retrieval system supports conditional requests, an ETag revalidation can confirm the selected representation without retransmitting its body. Preserve the earlier body and a separate revalidation event. A response with no new body is not an empty pricing page, and a validator is not a replacement for the saved evidence.
Test the collection shape before building a monitor
Try one permitted public URL with the free URL-to-Markdown tool and inspect its actual response fields. Its documented cache can retain successful results for up to 24 hours, so it is suitable for this collection prototype, not proof of a fresh pricing check.
The following Python script uses an installed curl executable and writes a new directory for each attempt. The timeout is a sample request budget. It saves the response and headers before judging success; valid text still requires review before it can become a comparable offer.
import hashlib, json, subprocess, urllib.parse
from datetime import datetime, timezone
from pathlib import Path
from uuid import uuid4
url = "https://anycrawler.com/"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urllib.parse.urlencode({"url": url})
started = datetime.now(timezone.utc).isoformat()
bundle = Path("captures") / uuid4().hex
bundle.mkdir(parents=True, exist_ok=False)
result = subprocess.run([
"curl", "--silent", "--show-error",
"--max-time", "60", "--dump-header", str(bundle / "headers.txt"),
"--output", str(bundle / "response.json"),
"--write-out", "%{http_code}", endpoint,
], capture_output=True, text=True)
http_status = int(result.stdout) if result.stdout.isdigit() else None
if result.returncode or http_status != 200:
failure = {"http_status": http_status, "exit_code": result.returncode,
"error": result.stderr, "collector_started_at": started}
(bundle / "failure.json").write_text(json.dumps(failure), encoding="utf-8")
raise RuntimeError(f"Capture failed; inspect artifacts in {bundle}")
raw = (bundle / "response.json").read_bytes()
headers = {}
for line in (bundle / "headers.txt").read_text().splitlines():
if ":" in line:
key, value = line.split(":", 1)
headers[key.lower()] = value.strip()
data = json.loads(raw)
markdown = data.get("results", {}).get("markdown")
valid = data.get("status_code") == 200 and isinstance(markdown, str) and bool(markdown.strip())
if isinstance(markdown, str):
(bundle / "page.md").write_bytes(markdown.encode("utf-8"))
record = {
"requested_url": url,
"collector_started_at": started,
"collector_finished_at": datetime.now(timezone.utc).isoformat(),
"http_status": http_status,
"target_status": data.get("status_code"),
"final_url": data.get("final_url"),
"api_timestamp": data.get("timestamp"),
"credits_used": data.get("credits_used"),
"response_sha256": hashlib.sha256(raw).hexdigest(),
"markdown_sha256": hashlib.sha256(markdown.encode()).hexdigest() if isinstance(markdown, str) else None,
"response_cache_control": headers.get("cache-control"),
"response_age": headers.get("age"),
"text_state": "review_required" if valid else "invalid",
"screenshot": None,
"offer": None,
"comparison_state": "not_comparable",
}
(bundle / "record.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
print(json.dumps({"bundle": str(bundle), **record}, indent=2))
if not valid:
raise RuntimeError("Saved response does not satisfy the text contract")
This collector was executed against AnyCrawler's public homepage. The endpoint and target status were both 200, and the Markdown included the pricing section. The response's API timestamp preceded the collector's completion time, consistent with the returned cache information. Neither final_url nor credits_used was present, so both remain null. The test establishes that this response can be retained and inspected; it does not establish a fresh quote, a verified currency, or a successful screenshot capture.
Open the saved page.md and identify the complete plan block, including billing and eligibility wording. Add the extracted offer only after checking those fields. Attach a separately verified screenshot if your acceptance rule needs one. A missing value stays unknown; the script deliberately leaves offer and screenshot empty and the comparison state blocked.
Decide what a change record is allowed to say
Compare the current observation with the last accepted baseline for the same offer key. Compare extracted fields for business meaning and use text or image differences to locate what a reviewer should inspect. A whole-page hash can change because of navigation or unrelated copy; it cannot identify a price increase by itself.
| Observation | Classification | Next action |
|---|---|---|
| Request failed or target card is missing | Capture failure | Keep the baseline; preserve failure and investigate |
| Currency, interval, quantity, or audience differs | Not comparable | Establish the correct state or create a separate offer key |
| Required fields agree and capture meets freshness rules | No observed offer change | Record the check without claiming universal stability |
| Comparable amount or terms differ | Candidate change | Review old and new text, conditions, and required images |
| Reviewer confirms the supported difference | Accepted change | Save the decision and advance that offer's baseline |
A useful candidate notification names the plan and changed field, carries the two observation identifiers, and states unresolved conditions. It should not say “the company raised prices” when the evidence only shows that an anonymous page displayed a different amount. If the old record lacks the billing interval, there may be no defensible before-and-after price comparison at all.
Do not promote an incomplete observation simply because a request succeeded. Keep retries bounded and separate operational failures from commercial changes. When the source removes a plan card, record “plan not found in this capture” until a complete read and review support a stronger conclusion. A failed extraction is not evidence that the offer was discontinued.
How much evidence should an alert require?
The remaining design choice is the cost of acting too soon versus the cost of waiting for review. An internal research queue may accept an explicitly unverified candidate. A customer-facing comparison needs a stronger rule about freshness, matching terms, and which source material a reviewer can inspect. The same polling interval and approval rule need not serve both uses.
Begin with one public page and one precisely defined offer. Inspect what the collector misses before deciding how often to run it or which changes should trigger a notification. If the source continually varies by visitor and you cannot record or reproduce that state, the responsible output may remain “different observed offers” rather than a single price history. Which decision will consume the alert, and what missing evidence would change that decision?
Frequently asked questions
Is a screenshot enough to prove a pricing change?
A screenshot supports a claim about visible content in the captured state. It may omit a footnote, selected billing control, region, or material below the captured area. Preserve its URL, timestamp, framing, and related text record. Compare it with an earlier capture of the same offer conditions, and review any disagreement. It does not by itself prove that every visitor could buy the offer or that the supplier changed its general pricing policy.
How should monthly and annual prices be compared?
Keep the billed interval separate from the displayed unit. A card can show a monthly equivalent while requiring an annual payment or commitment. Preserve that wording and give each billing arrangement its own offer key. Only compare successive records within the same arrangement, quantity, currency, and eligibility. If an earlier record lacks those details, mark it incomparable instead of calculating a percentage change from the two visible numbers.
Can cached content be used for pricing monitoring?
It can be used when its known age meets the consuming task's freshness requirement. Record the original observation time separately from your current request time and preserve any available cache metadata. If age is unknown, say so. The free collection example allows cached responses and therefore does not prove a new pricing observation on every run. A production workflow needs explicit freshness rules and a failure state when they cannot be satisfied.
Why is HTTP 200 insufficient for accepting a price record?
The returned page can omit the intended card, contain only a JavaScript shell, show a different regional offer, or lose qualifications during extraction. Check the target plan, amount, currency, unit, billed interval, and relevant conditions against the saved material. A successful request with incomplete evidence remains a capture requiring review. Keep the accepted baseline intact until the new observation satisfies the same comparison contract.
Does AnyCrawler provide the complete monitoring workflow?
The workflow here combines page-access capabilities with application-owned logic. Fetch and Render obtain page content; Screenshot provides a visual artifact, while Search helps discover sources when needed. Your application supplies scheduling, evidence retention, offer extraction, comparisons, review decisions, and notifications. Tasks that require selecting controls, signing in, or submitting forms need an authorized interactive browser workflow. The tested free collector demonstrates storage of one public response, not a complete monitoring service.






