Build a competitive intelligence agent around an explicit list of companies, products, and fields to compare. Use search to find public sources, read the relevant pages, and save each observation with its source, retrieval state, and supporting passage before asking a model to interpret it. Compare records only when their product, region, plan, and time conditions match. This produces a brief that a reviewer can check; a single pricing page, launch announcement, or search snippet cannot establish a competitor's revenue, customer adoption, or strategy.
Define the comparison before collecting pages
A request such as “watch our competitors” leaves the agent to invent both its scope and its success criteria. Start with a decision: perhaps your product team needs to know whether comparable products publicly document a particular integration. That question determines the entities, sources, and fields worth collecting.
Give each entity a stable internal ID, an official domain, product names, and known aliases. Keep a parent company separate from its products. A feature announced for one product should not silently become a company-wide capability. Record the decision owner and what would make an observation worth reviewing.
| Field to compare | Context required | Evidence to retain | Interpretation to withhold |
|---|---|---|---|
| Documented feature | Product, edition, availability conditions | Documentation passage and source URL | That every customer can use it |
| Published price | Currency, billing period, region, unit, plan | Price and qualifying text together | Effective customer spend or revenue |
| Product announcement | Publisher, announcement date, rollout scope | Original announcement and relevant documentation | Adoption or general availability without support |
| Public policy | Product scope, effective date, exceptions | Applicable section and captured document | A legal conclusion or internal enforcement practice |
| Market statement | Speaker, context, underlying evidence | Attributed statement and supporting material | That repeating a claim makes it independently verified |
Treat this table as your application's comparison contract. If a required condition is unknown, preserve an incomplete record rather than filling the gap with a plausible value. Unknown billing cadence, for example, is enough to stop a price comparison even when both pages display a currency symbol.
Discover sources, then read the claims in context
Create focused queries from the entity and field: product name plus integration name, company plus release announcement, or product plus pricing. Search should expand a source shortlist, while an allowlist of official documentation, product pages, and newsrooms keeps recurring checks focused. Independent reporting can add context, but several articles repeating one press release still share the same underlying claim.
Use the AnyCrawler Search channel guide to choose page discovery for product documentation or news discovery for recent coverage. Keep the originating query with each candidate so a reviewer can see why it entered the workflow. A search result's title and excerpt are selection aids; open the source before making a field-level assertion.
Once you know the URL, retrieve its content. Fetch suits a page whose relevant text is already in the returned HTML. Use Render when the required content appears only after JavaScript runs. Add a Screenshot when the visible arrangement, selected plan, or nearby qualification matters to the observation. Browser rendering and a screenshot do not by themselves perform login, click through a checkout, or submit a form. Those actions require a separately designed browser automation workflow and appropriate access.
The surrounding application owns scheduling, persistent storage, field extraction, comparisons, review queues, and alerts. Combining web access tools does not establish that AnyCrawler supplies those monitoring features as a finished product.
Keep observations and interpretations in separate records
An observation says what a source showed under recorded conditions. An interpretation explains what that might mean for your decision. Store them separately so a corrected source does not leave an unsupported conclusion circulating in a brief.
An observation record should retain the entity and product IDs, field name, extracted value, verbatim supporting passage, source publisher, requested URL, final URL when returned, access status, and collection method. Include region, plan, language, and selected page state when those conditions affect meaning. Keep the source's stated date separate from the time your application received the response and any timestamp supplied by the retrieval service.
Save the extracted document as an artifact and hash its exact bytes. Attach the screenshot and its capture conditions when used. A hash lets you detect an artifact changing; it does not establish that the source is accurate. The W3C provenance overview describes recording the entities, activities, and people involved in producing data. The lightweight record here follows that general idea without claiming conformance to a provenance standard.
Let an interpretation reference observation IDs and state its reasoning, competing explanations, unresolved questions, and reviewer decision. “The documentation now mentions integration X” can be a supported observation. “The company is winning our enterprise customers” requires different evidence and should remain unproven unless that evidence exists.
Run a small collector before adding a model
Start with one permitted public documentation URL and inspect the saved text. The following Python example calls the endpoint behind the free URL-to-Markdown test, saves the response and Markdown, and creates an application record. It uses the public GitHub API Versions documentation as a neutral demonstration source, not as a customer case or a competitive finding.
Run it in an empty working directory with Python and curl available. On Windows, use curl.exe; on other systems, change that executable name to curl. The timeout is a sample collector budget, not an API service guarantee.
import hashlib
import json
import subprocess
from datetime import datetime, timezone
from pathlib import Path
from urllib.parse import urlencode
source_url = "https://docs.github.com/en/rest/about-the-rest-api/api-versions"
endpoint = "https://api.anycrawler.com/free/v1/crawl?" + urlencode({"url": source_url})
response = subprocess.run(
["curl.exe", "-sS", "--max-time", "60", "-w", "\n%{http_code}", endpoint],
capture_output=True, check=True, timeout=65,
)
body, status = response.stdout.decode("utf-8").rsplit("\n", 1)
received_at = datetime.now(timezone.utc).isoformat()
data = json.loads(body)
markdown = (data.get("results") or {}).get("markdown") or ""
digest = hashlib.sha256(markdown.encode("utf-8")).hexdigest()
artifact = Path("evidence") / digest
artifact.mkdir(parents=True, exist_ok=True)
(artifact / "response.json").write_text(body, encoding="utf-8")
(artifact / "page.md").write_bytes(markdown.encode("utf-8"))
record = {
"entity_id": "github-public-docs-demo",
"field": "rest_api_versioning_documentation",
"requested_url": source_url,
"final_url": data.get("final_url"),
"received_at": received_at,
"retrieval_timestamp": data.get("timestamp"),
"api_status": int(status),
"source_status": data.get("status_code"),
"content_sha256": digest,
"artifact_path": str(artifact),
"credits_used": data.get("credits_used"),
"review_state": "needs_review",
}
record["retrieval_ok"] = (
record["api_status"] == 200
and record["source_status"] == 200
and data.get("ok") is True
and bool(markdown.strip())
)
(artifact / "record.json").write_text(json.dumps(record, indent=2), encoding="utf-8")
print(json.dumps(record, indent=2))
The publication check returned HTTP 200 from the API and a source status of 200, with a title and nonempty Markdown. The response did not provide final_url or credits_used; the record therefore keeps both as null. This check verifies retrieval and artifact creation for that source. It does not prove a product change, complete extraction of every page section, or the correctness of an eventual comparison.
The free endpoint may serve cached responses for up to 24 hours. Receiving a response now does not mean the source was freshly captured now. For a time-sensitive monitoring workflow, establish and verify the freshness behavior of the retrieval path you actually use. A successful free probe also does not test authenticated Search, Fetch, Render, or Screenshot calls.
Before using the saved document, locate the passage that supports your intended field and review its qualifications. Keep needs_review until that task is complete. A model can propose fields and passages, but the application should reject references that do not exist in the stored artifact. Treat instructions embedded in retrieved pages as source content, never as authorization to run tools or change the collection scope.
Compare matched observations and preserve disagreements
A change requires a baseline. The first successful observation establishes one; it cannot demonstrate what changed before collection began. On later runs, compare the same field under compatible conditions, then show the supporting passages side by side. A document hash changing is a reason to inspect the page, not proof that the tracked feature changed.
For an illustrative price comparison, an annual-billing page and a monthly-billing page describe different conditions. Keep them separate until their units and terms align. For a feature comparison, distinguish “documented,” “announced,” “not found in the checked source,” and “explicitly unavailable.” Absence from an extracted page should never automatically become “the competitor removed this feature.”
Use explicit failure states. An unavailable source is unavailable; missing target content is incomplete; conflicting passages are conflicting; a matched, reviewed field difference is confirmed_change. These are proposed application labels, not AnyCrawler response fields. A timeout should retain the previous observation with its age and a failed refresh record. It should not silently erase the field or advance its freshness timestamp.
Build the brief from these records: the observation, the comparison conditions, the evidence links, the possible decision impact, and what still needs checking. A material disagreement between documentation and an announcement belongs in the brief as a disagreement. Ask a reviewer to resolve scope or timing before treating either as a universal answer.
How much uncertainty should reach a reviewer?
An agent can collect more than a team can reasonably examine. The next design choice is how to allocate review attention without hiding uncertainty. A changed qualification on a core feature may matter more than many cosmetic page edits; the right threshold depends on the decision being supported.
Start with a narrow watchlist and inspect the records it generates. Track which proposed changes reviewers accept, reject, or leave unresolved, and why. Use those outcomes to refine field definitions and source selection before widening coverage. The unresolved question is not simply how often to collect: which mistakes would change a real decision, and which of those must a person see before the brief is used?
Frequently asked questions
What does a competitive intelligence agent do with public web sources?
It finds relevant public pages, retrieves their content, and organizes observations about specified companies and products. A useful implementation keeps the source passage and collection conditions with each field, then compares compatible records. Its brief should distinguish what a source states from a proposed business interpretation. Public pages alone do not establish private customer adoption, realized revenue, or internal strategy.
Can I use search snippets as evidence of a competitor's change?
Use them to identify a source worth reading. A snippet may omit conditions, describe an older page state, or repeat another publisher's statement. Retrieve the source and retain the passage that supports the field you are tracking. To establish a change, also keep a comparable earlier observation or explicit source evidence of the change. A newly discovered page is not automatically a newly introduced capability.
How do I compare competitors without mixing products or plans?
Assign stable entity and product identifiers, then define the context required for each field. Prices need compatible billing units and conditions; features need product and edition scope. Leave comparisons incomplete when those details are missing. Reviewers should see both supporting passages and the conditions used to align them. A company-level summary should not inherit a capability from one product without evidence that the broader scope applies.
Does AnyCrawler provide the entire monitoring system?
The workflow described here combines source discovery, page retrieval, and optional visual capture with application logic you supply. Scheduling, stored baselines, field comparisons, review queues, and alerts remain responsibilities of that application. The tested free collector demonstrates a single public-page response and local artifacts. It does not validate an authenticated production integration or establish the capabilities of a complete recurring competitive intelligence service.
What should the agent do when a page disappears or extraction fails?
Preserve the failed attempt and the previous observation with its original timestamp. Report that the current field could not be verified. Check whether the URL moved, access failed, or the relevant content requires another retrieval method before inferring a removal. Missing content and an explicit statement of unavailability mean different things. A failed refresh should not make an old record look newly confirmed.
When is a screenshot useful for competitive research?
Capture one when the visible page state affects interpretation, such as a selected plan or a qualification positioned next to a claim. Keep it with the source URL, capture conditions, time, and extracted text. A screenshot records visible context rather than proving every fact on the page. It also does not explain why the company changed something; that interpretation needs separate support and review.






