Build a news monitoring agent around an evidence record for each story: discover candidates with News Search, group related coverage, read selected sources, and preserve the text and any necessary screenshots before deciding what changed. Store the event time, article publication time, collection time, source URLs, and verification state separately. Your application owns the schedule, comparison logic, review policy, and notifications. AnyCrawler provides the web access steps you can combine. A useful first version produces a reviewable queue of new claims and corrections, with incomplete sources clearly marked, before it starts sending alerts.

Decide what deserves another look

Write the monitoring rule before choosing a model. “Follow this organization” leaves too much undefined. A more useful rule names the entity, the developments that matter, the sources worth checking, and the evidence needed before an item reaches a reader. A product launch monitor might care about an announced release date, an availability change, or a correction to an earlier announcement. A passing mention would remain in the candidate log.

Keep the queries and their version with each run. Changing a query can change the candidate pool; that does not establish a change in the world. Record a separate run_id and distinguish a successful search with no qualifying items from a search that failed. Both can produce an empty digest, but only the former means the monitor completed its checks.

Use the News Search request and response contract to build the discovery adapter. It returns story candidates and source context; selecting a result for reading is a separate decision. Start with a small set of precise queries, then inspect missed and irrelevant items before expanding them. Search ranking and snippets are triage inputs, not confirmation of the underlying report.

Keep documents separate from events

A URL identifies a candidate document. An event is the development that several documents may describe. Store both identities so a repeated URL can be recognized without erasing a new development in the same story.

For document deduplication, keep the requested URL, resolved URL when available, publisher-declared canonical URL, and content hash. Validate a canonical destination before using it to merge records. Hashes recognize identical saved text; they do not establish independent reporting or equivalent meaning. Navigation changes can also change a whole-page hash.

Event grouping needs a narrower question: do these reports describe the same entity, action, and occurrence? Similar headlines are a reason to compare the passages. They are not sufficient grounds to discard a report. Preserve each source's attribution and any disagreement inside the group.

Observed candidate Record action Delivery decision
Same document and unchanged relevant passage Reuse document identity; add the observation Usually suppress a repeat
Different URL carrying the same attributed wire report Keep both URLs with their shared reporting origin Do not count copies as independent confirmation
Same event with a new primary announcement Attach the announcement as another source Review whether it adds support or changes the claim
Same URL with a corrected date or claim Save a new document version and the affected passage Reopen the event for correction review
Similar headline about a different occurrence Keep separate event identities Evaluate each occurrence
Conflicting accounts of the same event Retain both claims and their supporting passages Mark the disagreement for review

An internal event_id should survive headline edits. A hash of the headline alone is a fragile event identity: a correction can change the key and make an old event look new. Assign a stable identifier after grouping, then version the accepted claim set beneath it.

Give every timestamp a job

Use distinct fields for event_at, published_at, and observed_at. The first describes when the reported development occurred. The second comes from the source's publication information. The third records when your application collected the response. Keep a publisher update time separately when it exists.

Google's guidance on publication and update dates explicitly separates a page's dates from the events it describes. That distinction is useful in a monitoring ledger too. Compare visible bylines and structured metadata, retain the original value and timezone, and leave an ambiguous date unresolved. Do not convert an unqualified local time into a precise UTC timestamp by guessing a timezone.

Keep these fields alongside the dates:

  • source_id, event_id, and document_version, so a correction can be traced back to its source.
  • query, run_id, and discovered_at, so the discovery context remains visible.
  • Requested, final, and canonical URLs, each with its own meaning and nullable when unavailable.
  • Retrieval status, saved text, text hash, supporting passages, and any screenshot reference.
  • An evidence state such as candidate, collected_needs_review, supported, conflicting, or incomplete.
  • A review note naming the claim that was checked and the evidence that supports it.

These are application fields, not a promise that every API response contains them. A supported state should apply to a specific claim and source bundle. It should never mean that everything on the page is true.

Escalate from reading to rendering only for a reason

Read a selected public source through Fetch when its useful content is available in HTML. Check for the passage your monitor actually needs: the announced product name and availability statement, for example. A successful HTTP response with a navigation menu or a loading message cannot satisfy that check.

Use Render when the needed content depends on JavaScript and the page is accessible within your permitted workflow. Rendering does not supply missing access rights or turn login, clicks, forms, and other interactions into a supported reading task. Retain a blocked or incomplete result when the needed content cannot be collected.

Add a stored screenshot of the source page when a reviewer needs to see a correction notice, presentation context, or another visible state alongside the text. Save the capture time and settings with the image. If the screenshot and extracted passage disagree, investigate the different page states before accepting the bundle. A screenshot preserves what was visible in that capture; it does not authenticate the publisher or prove all of the page's claims.

Run a collection probe before wiring the monitor

The following Python probe runs on Windows with Python and curl.exe available. It uses an archival NASA article about Webb's first deep field. It deliberately tests collection and record handling rather than discovery of breaking news. It calls the public free endpoint, checks for a source-specific phrase, and leaves the result awaiting review. The 60-second curl timeout and 65-second process timeout are local example settings, not service guarantees.

import hashlib
import json
import subprocess
from datetime import datetime, timezone
from urllib.parse import urlencode

source_url = (
    "https://www.nasa.gov/image-article/"
    "nasas-webb-delivers-deepest-infrared-image-of-universe-yet/"
)
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urlencode(
    {"url": source_url}
)
observed_at = datetime.now(timezone.utc).isoformat()
try:
    response = subprocess.run(
        ["curl.exe", "-sS", "-L", "--max-time", "60",
         "-w", "\n%{http_code}", endpoint],
        capture_output=True, check=True, timeout=65,
    )
    body, status = response.stdout.decode("utf-8").rsplit("\n", 1)
    http_status = int(status)
    data = json.loads(body)
except (OSError, subprocess.SubprocessError, ValueError) as error:
    http_status = None
    data = {"ok": False, "error": type(error).__name__}

result = data.get("results") or {}
markdown = result.get("markdown") or ""
has_expected_context = "SMACS 0723" in markdown
collected = (
    http_status == 200
    and data.get("ok") is True
    and data.get("status_code") == 200
    and has_expected_context
)
record = {
    "requested_url": source_url,
    "final_url": data.get("final_url"),
    "observed_at": observed_at,
    "retrieval_timestamp": data.get("timestamp"),
    "published_at": None,
    "event_at": None,
    "http_status": http_status,
    "source_status": data.get("status_code"),
    "title": result.get("title"),
    "content_sha256": hashlib.sha256(markdown.encode("utf-8")).hexdigest(),
    "credits_used": data.get("credits_used"),
    "evidence_state": "collected_needs_review" if collected else "incomplete",
    "error": data.get("error"),
}
print(json.dumps(record, indent=2, ensure_ascii=False))

The publication-day check returned the expected title and passage with a successful source status. The extracted Markdown also contained navigation, so this probe does not establish clean article boundaries. The response omitted final_url and credits_used; the record preserves them as null. Here observed_at marks the start of the collection attempt, while retrieval_timestamp preserves the time supplied by the service. An earlier incorrect source path returned a source HTTP error and was corrected rather than repeatedly retried.

Try a source you intend to monitor in the free URL-to-Markdown tool, inspect the relevant passage, and adapt the content check before building the recurring job. The free endpoint can serve a cached response for up to 24 hours. It is suitable for this shape check, not proof of a fresh recrawl or a latency-sensitive monitoring contract. Authenticated News Search, Fetch, Render, and Screenshot calls require their own integration tests.

The probe prints a compact record and does not persist the full evidence bundle. In an application, save the response and Markdown with the record before marking collection complete. Then isolate the relevant passage and compare it with the source. A phrase match can detect missing context; it cannot decide whether an article supports the claim your digest will make.

Make repeated runs and corrections reviewable

Persist the candidate and its retrieval outcome before asking a model for a summary. Give the summarizer the relevant source passages and stable source IDs. Treat retrieved text as material to analyze; instructions embedded in a page must not change the monitor's tool permissions or delivery rules. Require the output to identify which passage supports each proposed claim, and reject unsupported additions.

Store accepted event versions separately from delivery attempts. A delivery key can combine the stable event ID, accepted version, destination, and policy version. Check it when queuing a notification so a repeated search does not enqueue the same accepted item again. Delivery retries and acknowledgment handling still belong to your application; a key alone does not guarantee exactly-once delivery.

Exercise the monitor with an unchanged source, a correction at the same URL, a syndicated copy, conflicting reports, and a failed retrieval. Confirm that each produces the intended state. A failed page should retain its error and leave other sources available. A correction should reopen the affected claim, while an unchanged document should preserve history without creating a new alert.

Run health belongs in the output as well. Report which queries completed, which sources remained unreadable, and whether the review queue contains unresolved items. This lets a reader distinguish a quiet news cycle from a monitor that lost access to its evidence.

How much uncertainty can an alert carry?

The unresolved design choice is how quickly a reader needs to know compared with the cost of correcting a premature alert. A routine digest can wait for a readable primary source. A time-sensitive workflow may choose to show a provisional item sooner, provided its source, unresolved claim, and review state remain visible.

Decide what additional evidence would promote, revise, or withdraw that item. More copies of the same report do not necessarily add support; a correction from the original publisher may matter more. Before enabling unattended delivery, define the answer to this question: which unresolved evidence states may reach this audience, and who owns the follow-up when they change?

Frequently asked questions

Does AnyCrawler schedule news monitoring or send alerts?

The architecture here uses AnyCrawler for discovery, page reading, rendering, and optional visual capture. Your application supplies the recurring trigger, saved state, event comparison, review policy, and delivery integration. Build and verify a single collection cycle before adding recurrence. Keep delivery outcomes separate from evidence status so a notification retry does not become a new story or make an unsupported claim appear reviewed.

How should a news monitoring agent handle syndicated articles?

Keep the individual documents and group copies that share a reporting origin. Similar titles can help find candidates for comparison, but inspect attribution and the relevant passages before merging them. Preserve original reporting and conflicting claims within the event record. The number of websites carrying a story is not the number of independent confirmations, and the earliest visible timestamp does not by itself prove authorship.

Should every news page be rendered and screenshotted?

Choose the access step from the evidence needed. Fetch can be sufficient when HTML contains the relevant article passage. Render is useful when accessible JavaScript content is necessary to complete that passage. A screenshot adds visible context when presentation or a notice matters to review. Check the resulting content in every case, and retain an incomplete state if the required evidence remains inaccessible or inconsistent.

Which date should determine whether a story is new?

Define that in the monitoring policy and retain the dates needed to explain the decision. Publication time describes the source document, event time describes the development, and observation time describes your collection. A newly discovered old article may be new to the monitor without reporting a new event. Preserve missing or ambiguous dates explicitly, and treat a correction as a new document version when its relevant claim changes.

What should happen when one source cannot be read?

Save the failure with its source URL, observation time, and retrieval outcome. Continue evaluating other accessible sources while keeping the missing evidence visible. A search snippet cannot substitute for an unread article when the claim requires its text. Correct an invalid URL or request before retrying, use bounded retries for transient failures, and let the review policy decide whether the remaining evidence is sufficient for a provisional item.