A screenshot belongs in an AI agent's evidence record when the claim depends on what a person could see, not only on text the agent could extract. Layout, a selected option, a disclosure beside a price, an error banner, and a rendered chart can all change the meaning of otherwise correct Markdown. Save text for search and reasoning, save a screenshot for visible state, and save both when a reviewer must connect the claim to the page. The practical test is simple: ask what would be lost if either artifact disappeared.

The decision starts with what must be proven

Choose the evidence format from the claim, not from the convenience of the capture tool. If the agent needs to quote a policy paragraph, compare headings, or index a document, machine-readable text is usually the primary record. If the claim is that a warning appeared above a button, a plan was visually highlighted, or a chart rendered with a missing series, pixels carry information that text extraction can flatten or omit.

Some tasks need both. A screenshot helps a reviewer inspect placement and state, while Markdown makes the same record searchable, quotable, and accessible to downstream code. The W3C explanation of images of text draws a useful boundary: text in an image is not the same as text that software can determine and present in different ways. A screenshot should therefore supplement structured content, not silently replace it.

Evidence choice Use it when It can answer It cannot establish by itself
Text only Meaning is carried by headings, paragraphs, links, metadata, or values What the page says and how the document is structured Exact placement, styling, overlap, or visible UI state
Screenshot only The task is a bounded visual observation and no semantic extraction is required What one rendered frame looked like Complete page meaning, hidden content, source integrity, or legal compliance
Both A human must review a claim that depends on wording and presentation What the page said, where it appeared, and what state was visible Events before or after capture, intent, or facts outside the captured page

That table is also a scope guard. "Both" is not automatically better; it creates more storage, more retention decisions, and another artifact that must be verified. Use it when losing either the semantic or visual side would make the record materially weaker.

A screenshot records a rendered frame, not the whole web page

The phrase "page screenshot" hides several capture contracts. The current W3C WebDriver working draft defines commands for a top-level browsing-context screenshot and an element screenshot. Browser tools may add other modes: Playwright documents viewport, full-page, buffer, and element captures. Those modes are not interchangeable.

A viewport capture can prove that an element was visible inside a particular frame. It says nothing about content below the fold. A full-page image widens coverage, but sticky headers, lazy loading, animations, and very long pages can make the result different from a single human viewport. An element capture is precise, yet it may remove the surrounding label, price qualifier, or warning that gives the element meaning.

Rendering conditions matter as well. Playwright's visual comparison guidance warns that screenshots can vary with operating system, browser version, settings, hardware, and headless mode. For evidence workflows, this means a pixel difference is an observation to investigate, not an automatic finding. Keep capture mode and environment stable when comparing images, and route material changes to review.

Build one evidence record around both artifacts

Do not leave the image URL in one system and the extracted text in another with no shared identifier. Store a small manifest that binds the requested URL, final URL, capture result, text result, and review state. Use an RFC 3339 timestamp so different services can exchange an unambiguous capture time.

const endpoint = "https://api.anycrawler.com/v1/crawl/screenshot";
const requestedUrl = "https://example.com/";

const response = await fetch(endpoint, {
  method: "POST",
  headers: {
    authorization: `Bearer ${process.env.ANYCRAWLER_API_KEY ?? ""}`,
    "content-type": "application/json",
  },
  body: JSON.stringify({ url: requestedUrl, full_page: false }),
});

const payload = await response.json();

if (!response.ok || !payload.ok) {
  throw new Error(JSON.stringify({
    httpStatus: response.status,
    errorCode: payload.error_code,
    retryable: payload.retryable,
    creditsUsed: payload.credits_used ?? 0,
  }));
}

const evidence = {
  requestedUrl,
  finalUrl: payload.final_url,
  capturedAt: new Date().toISOString(),
  statusCode: payload.status_code,
  creditsUsed: payload.credits_used,
  screenshot: {
    url: payload.results.snapshot_url,
    key: payload.results.snapshot_key,
    bytes: payload.results.snapshot_bytes,
    imageType: payload.results.snapshot_image_type,
  },
  review: { state: "pending", notes: [] as string[] },
};

console.log(JSON.stringify(evidence, null, 2));

The failure branch is part of the example, not an afterthought. Authentication, credit, rate-limit, browser, and target-page failures should never produce a record marked as captured. Keep retryable with the error, retry only finite transient failures, and send terminal errors for configuration or human review.

Use the AnyCrawler Screenshot API when a public page needs a stored image artifact. If the workflow also needs Markdown, run the lightest complete read path separately: start with Fetch when useful content is in the returned HTML, and move to Render only when JavaScript is required.

Capture after the target state becomes reviewable

Timing is part of the evidence contract. A technically successful image can still be useless if it contains a loading skeleton, cookie panel, empty chart, or animation midpoint. Before capture, define the visible assertion in concrete terms: the price card is present, the warning text is adjacent to the action, the selected tab is active, or the chart legend and series are both visible.

For public on-demand capture, keep the checklist narrow:

  1. Record the requested URL and accept the final URL only after redirect handling completes.
  2. Decide whether the claim needs the viewport, the whole scrollable page, or a specific element with context.
  3. Wait for the assertion's required state, not merely for an HTTP response or a generic page-load event.
  4. Save the capture time, result status, image key, image type, byte count, and request identifier exposed by the service.
  5. Read the stored artifact back before treating it as durable evidence.

This workflow does not imply login interaction, scheduled monitoring, visual diff alerts, or bypassing a target's controls. Those are separate capabilities with separate permissions and failure modes. An on-demand screenshot of a public page proves only the captured state returned by that request.

Validate the text and visual sides independently

A combined record is only as reliable as its weaker half. Validate each artifact before linking them.

For the text side, check the requested and final URLs, HTTP status, title or canonical URL, and the exact passage or field that supports the claim. Hash the normalized content if later integrity checks matter. Do not accept a successful request when the required sentence, table row, or value is missing.

For the visual side, open the image rather than trusting its storage URL. Confirm that it decodes, has non-zero dimensions, shows the expected page rather than an error screen, and includes enough surrounding context to interpret the target. Look for overlays, cropped qualifiers, blank lazy-loaded areas, and inconsistent viewport framing.

Finally, validate the join:

  • Both artifacts refer to the same requested URL and compatible final URL.
  • Their capture times are close enough for the page's rate of change.
  • The manifest distinguishes observed fields from reviewer conclusions.
  • A reviewer can trace the claim to a text location and a visible region.
  • Failure or uncertainty remains visible instead of being converted into a pass.

The result is not a universal truth about the page. It is a bounded record of what two capture paths returned under documented conditions.

Route failures before they contaminate the evidence store

Treat screenshot evidence as a small state machine: requested -> captured -> artifact verified -> paired -> reviewed. A request can enter failed from any state, but it should never skip artifact verification. This prevents a JSON success envelope, broken image, or mismatched final URL from becoming a trusted record.

Use the error class to choose the next action. Missing or invalid credentials require configuration, not repeated requests. Insufficient credits require an account decision. Rate limits, concurrency conflicts, and retryable upstream failures may justify bounded backoff. A target timeout or incomplete visual state may require a different capture scope or a human decision. Keep the original error and request identifier even when a later attempt succeeds.

For the broader tool boundary, the Search, Fetch, Render, or Screenshot decision guide explains how to choose a web-access path before building this evidence layer. Screenshot capture belongs to the read-and-prove path; login, clicking, form submission, and other side effects belong to browser automation with explicit authorization.

When does screenshot evidence become stale?

There is no universal retention or recapture interval. A static documentation page, a personalized price, and a live incident banner change at different rates. The open decision is therefore not "How often should every agent recapture pages?" but "Which change would make this claim misleading, and how would the workflow detect it?"

Teams still need to choose whether time-based recapture, an upstream event, a content-hash change, or a human request should reopen the evidence record. They also need rules for sensitive visual data, retention, and access. A screenshot can strengthen review, but it does not decide those policies and does not by itself prove authenticity, ownership, legal compliance, or what happened outside the captured frame.

Frequently asked questions

Is a screenshot enough evidence for an AI agent?

Usually not when the agent must reason about or quote page content. A screenshot preserves visible presentation, but it does not reliably expose headings, links, metadata, hidden content, or exact text to downstream systems. Pair it with structured extraction when wording and layout both matter. A screenshot can be sufficient for a narrowly defined visual observation, provided the record also keeps the URL, time, capture scope, and verification result.

When should an agent save both Markdown and a screenshot?

Save both when removing either artifact would weaken the claim. Examples include a disclosure whose placement matters, a selected plan with nearby qualifiers, an error message tied to a particular control, or a rendered chart that needs its labels preserved as text. Markdown supports search and quotation; the screenshot supports human inspection of visible state. Link them through one manifest instead of treating them as unrelated outputs.

Does a full-page screenshot prove that every part of the page loaded?

No. Full-page describes capture scope, not content completeness. Lazy-loaded sections may still be blank, overlays can hide content, animations can be caught mid-state, and sticky elements may repeat or cover regions. Define required visual assertions before capture, then inspect the stored image for those assertions. When exact wording matters, independently verify the extracted text rather than inferring it from pixels.

Can screenshot evidence prove legal or regulatory compliance?

Not by itself. A screenshot can document a visible state for review, but compliance conclusions depend on the applicable rule, jurisdiction, process, timing, and evidence chain. The image cannot prove what happened before or after capture, whether the page was personalized, or whether the record is complete. Treat it as one bounded artifact and route legal conclusions to qualified review with the supporting text and metadata.

What metadata should accompany a screenshot?

At minimum, keep the requested URL, final URL, capture timestamp, capture mode, target status, request identifier, image key or URL, image type, byte count, and verification state. Add viewport or environment details when visual comparison depends on them. If the screenshot is paired with text, record the text artifact identifier or hash and the passage supporting the claim. Keep reviewer notes separate from fields returned by the capture service.

Should an agent automatically retry failed screenshot requests?

Only when the failure is explicitly transient and the retry budget is bounded. Rate limits, concurrency conflicts, or retryable upstream failures may justify backoff. Missing credentials, insufficient credits, unsupported inputs, and policy decisions need configuration or human action instead. Preserve the first error and request identifier, and never mark evidence as captured until the stored image has been read back and visually validated.