Cache web access only when the saved representation still matches the request, meets the task's freshness requirement, and contains the evidence the answer needs. A successful request or an unexpired cache entry does not establish all three. For a research agent, keep the time you received a response separate from the time its content was obtained or validated, and preserve the original snapshot when refreshing it. If age or provenance is unknown, treat the result as unverified for current claims. This lets you reuse useful material without silently presenting historical content as a new observation.
A new response can contain yesterday's evidence
Consider a public crawl request for Example Domain. The response can arrive now, contain the expected title and Markdown, and still come from a cache. In the diagnostic run behind this example, both the API request and target status were successful, while the response included a cache hit and an earlier crawl timestamp. That is a useful extraction result. Whether it is useful evidence depends on the question.
For “what does this example document say?”, the stored text may be sufficient. For “has this page changed since the last check?”, the same response cannot establish that a new source observation occurred. Receiving identical Markdown twice also cannot distinguish an unchanged source from a reused extraction.
This is the central web scraping cache problem: delivery time, content age, and suitability for a particular claim are separate properties. Keep each visible in your data model instead of allowing the latest request time to stand in for all of them.
Set the freshness requirement before choosing a TTL
Start with the consequence of being wrong. The following classes are an application design aid, not service defaults or measured refresh intervals.
| Reader task | Reuse condition | When the condition fails |
|---|---|---|
| Explain a named document version | The stored version and required passage match | Obtain the specified version; do not silently substitute the latest |
| Answer a question about current documentation | A trusted observation or validation time fits the task's age budget | Validate or retrieve again before calling it current |
| Check a changing announcement or availability field | The workflow has evidence recent enough for that specific check | Return pending or unavailable if a current observation cannot be obtained |
| Reproduce an earlier research answer | The original snapshot and citation remain identifiable | Recover the historical evidence; a new page is not a replacement |
| Offer background context during an outage | The caller explicitly accepts labeled historical material | Keep the failure visible and omit unsupported current claims |
A TTL controls when your application stops reusing an entry without another check. Retention controls how long you keep the entry. A document can be too old for a current answer yet still need to remain in the evidence archive. Deleting it at TTL expiry makes later explanation of an earlier answer harder.
Assign an age budget to the task or field, then record the policy version with the decision. A change to that budget should affect later reuse decisions without rewriting when the original observation happened. Avoid choosing a shorter interval simply because a request was quick; request speed says nothing about how often the source changes.
Match the representation, then evaluate its age
At the HTTP layer, cache selection depends on the request and relevant representation variants, including fields named by Vary. An unqualified no-cache response requires successful validation before reuse; no-store prevents storage, and must-revalidate restricts reuse after staleness. These are protocol rules, not a guarantee that the page's claims are accurate. See the HTTP caching standard when implementing that layer.
Your extraction cache needs a separate identity. Include the requested URL and every supported option that changes the output: fetch versus render, extraction format, language context, and the version of your own extractor or transformation. Partition any authorized private context so one caller cannot receive another caller's representation. Do not indiscriminately remove query parameters; a parameter may select a document version or a different result.
Save these fields with each accepted result:
received_at: when your application received this response.provider_timestamp: the timestamp returned by the service, with its documented meaning.source_observed_atorvalidated_at: only when the retrieval path provides trustworthy evidence for that event.requested_url,final_url, and extraction options: preserve missing fields as missing.content_hash, snapshot location, and content checks: identify the exact material used.
Never copy received_at into source_observed_at merely to make an entry pass the age check. A provider timestamp can also need interpretation: it may describe a cached response, while an HTTP Age header can describe an intermediary's response rather than the original page observation.
Revalidation updates a decision, not the historical snapshot
When a direct HTTP client has an origin validator, it can send If-None-Match with the stored ETag. A valid 304 Not Modified response lets it reuse the associated representation without downloading another body. The MDN caching guide explains that conditional request flow. If the service you call does not expose origin validators or conditional requests, do not invent those fields in its API payload.
Record a successful validation as a separate event referencing the stored content hash. Keep the original capture time intact. A page composed from several resources still needs an extraction check: validation of a particular HTTP representation does not by itself establish that every client-loaded field was observed or that every factual claim remains correct.
For JavaScript-dependent content, review the documented AnyCrawler render cache controls. The Render API exposes accept_cache, which defaults to false. That is a product-specific control; it does not define your application's evidence policy, and the free endpoint has its own cache behavior. Search discovers candidates, Fetch reads available HTML, Render executes a page for extraction, and Screenshot preserves visible state. None of these roles supplies your application's complete scheduling, archival, or freshness decision process.
Inspect a real cached response before writing the policy
The free URL-to-Markdown playground lets you inspect a public page without an API key. Its documentation says successful responses may be cached for up to 24 hours. That makes it suitable for checking response shape and for demonstrating why a fresh network request is not necessarily a fresh crawl.
Run this in an empty working directory with curl installed. On Windows PowerShell, use curl.exe for the executable. The files preserve the observed headers and body for inspection; use a new directory for each run if you want to retain earlier captures.
curl --silent --show-error --fail --max-time 60 \
--dump-header response-headers.txt \
--output response.json \
'https://api.anycrawler.com/free/v1/crawl?url=https%3A%2F%2Fexample.com'
The checked request returned HTTP 200, target status_code: 200, and the title Example Domain. Its headers contained CF-Cache-Status: HIT and Age: 76302; the JSON timestamp was 2026-09-11T03:49:33.385Z, whereas the response Date header was 2026-09-12T01:01:15Z. Those values describe this observation only. The result did not include final_url or credits_used, so the record leaves them unavailable. The command's 60-second timeout is a local stop condition, not a service limit or latency claim.
Inspect the expected paragraph as well as the statuses. For a real task, replace that paragraph check with a required heading, field, or passage. An error page can be long; a valid example page can be short. A character-count threshold alone is a poor acceptance rule.
The next function models your application's decision once it has a trusted observation age. It deliberately does not turn the free response's timestamp into that age automatically. Here, max_age is chosen by the caller in the same units as age; the example is not an HTTP cache implementation.
def reuse_decision(*, identity_matches, content_valid, age, max_age):
if not identity_matches:
return "retrieve_matching_representation"
if not content_valid:
return "reject_content"
if max_age < 0:
raise ValueError("max_age must be nonnegative")
if age is None or age < 0:
return "freshness_unknown"
if age > max_age:
return "validate_or_retrieve"
return "reuse"
The function was checked with matching and mismatching identities, valid and invalid content, missing and negative ages, an exact age boundary, and an expired entry. Missing age produces freshness_unknown; it does not earn a new lifetime just because the response was received successfully. In production, validate numeric inputs and timestamp formats at the boundary before calling this function.
Keep refresh failures out of the success path
When validation or retrieval fails, retain the last accepted snapshot and create a failure event. Do not replace a useful document with a timeout message, an access-denied page, or an empty extraction. Keep retry limits and refresh status separate from the stored evidence's age.
HTTP also supports availability-oriented stale delivery. stale-while-revalidate can serve an older response while validation runs; stale-if-error can permit bounded stale fallback when specified failures occur. The stale response extensions describe those mechanisms. They do not turn a fallback into newly verified evidence.
Make the application choose an explicit outcome: wait for a refresh, return unavailable for the current claim, or return historical context with its observation time and refresh failure attached. If the caller requested a current check, a background refresh that has merely started cannot satisfy that request yet.
For concurrent refreshes, associate the result with its request identity and observation event before updating the reusable entry. A slower, older attempt should not silently overwrite a newer accepted observation. Preserve both events if they matter to the audit trail, and let the acceptance rule decide which entry subsequent requests can use.
Which questions deserve strict freshness?
The difficult choice is often whether a temporarily incomplete answer is preferable to a complete answer backed by older material. That decision can differ even within one page: a stable explanation may remain useful while a changing availability field needs another observation. Measure the fields whose changes would alter the answer, and review rejected or outdated observations before changing the age budget.
Start by inspecting one response's headers, provider timestamp, and required content. Then write down what your application should do when any of those cannot support the requested freshness. If the product cannot tell a reader which observations support “current,” should it make that claim at all—or show a dated answer until verification succeeds?
Frequently asked questions
Does an HTTP 200 response prove a crawl is fresh?
No. It indicates a successful HTTP response, which can be delivered from a cache. An extraction API can also report a successful target status inside a cached JSON body. Check the relevant timestamps and cache indicators, then decide whether they establish an observation recent enough for the task. Record unknown age explicitly when the service does not provide enough information, and still verify that the required content is present.
What TTL should I use for a web scraping cache?
Choose it from the task's tolerated evidence age and the consequence of an outdated answer. A named historical document, current documentation, and a changing availability field need different policies. Keep the budget configurable and record which policy accepted each result. A TTL is a reuse limit, not proof that the source cannot change during that interval. Retain historical snapshots separately when later review requires them.
Can I reset the timestamp when I reuse cached Markdown?
Record a new received or served time if that event matters, but preserve the original observation time. Otherwise an old extraction can appear indefinitely new after repeated reuse. If you successfully validate the associated representation, record that validation as a separate event linked to its content hash. When the upstream time is missing or ambiguous, do not manufacture it from your application's clock; mark freshness unknown for current claims.
Does a content hash tell me whether a page is current?
A hash identifies the exact content you stored and helps connect an answer to its snapshot. It does not establish when that content was observed or whether a newer source version exists. Identical hashes from repeated cached responses therefore do not prove repeated source checks. Combine the hash with request identity, reliable timing information, and content acceptance checks before deciding whether the result can support a current answer.
Should an agent use stale content when a refresh fails?
Only when the caller's policy permits historical context and the answer makes its age and failed refresh clear. Keep the last accepted snapshot, but do not relabel it as a successful current check. For tasks that require a new observation, return a pending or unavailable state until verification succeeds. Background refresh and stale-on-error behavior improve availability; your application still owns the decision about which claims that evidence can support.






