Web scraping retries and timeouts should preserve useful results while placing a firm limit on failed work. Give each URL its own attempt history, distinguish a temporary transport failure from missing content, and stop when the time or request budget runs out. A successful HTTP response is only one checkpoint: the extracted document still has to contain the evidence your task needs. The practical pattern is a bounded retry loop around each permitted read, followed by a batch result that keeps successful documents and explicitly lists unresolved inputs.
Classify the failure before choosing another request
A web access pipeline has several places to fail. The client can lose its connection to the API; the API can fail to retrieve the target; or the returned text can be unusable for the task. Store the outer HTTP status separately from a target status inside the response. Otherwise a gateway response can hide the difference between a missing article and an unavailable service.
Use the following table as an application policy to adapt to the provider's contract. The state names are proposed labels for your own records.
| Observation | Record state | Next action |
|---|---|---|
| Connection reset or attempt timeout during a permitted read | Retry candidate | Check remaining budget and repeat only when request semantics allow it |
| Gateway throttles the caller | Retry candidate | Respect its delay advice and reduce pressure at the relevant account or host scope |
| Temporary upstream failure | Retry candidate | Apply bounded backoff; inspect provider error details before assuming every server error is temporary |
| Authentication, credit, permission, or invalid-input problem | Permanent for this run | Correct the prerequisite before resubmitting |
| Gateway succeeds but target reports an error | Needs review | Inspect target status and explicit retryability; do not treat the outer success as content success |
| Document arrives but a required section is missing | Content invalid | Check URL and extraction rules; consider rendering only if JavaScript explains the gap |
| Attempt count or remaining time is exhausted | Exhausted | Keep the final reason and hand the unresolved item to the next workflow stage |
An HTTP 429 response specifically reports rate limiting and may include Retry-After, as defined in the HTTP rate-limit status specification. That is a reason to slow the relevant queue. Restarting the entire batch can make the limiting condition worse and needlessly reprocess good documents.
Provider meanings also matter. A server error caused by missing configuration needs operational attention even if the same status is commonly used for temporary overload. Preserve error_code, retryable, and the request identifier when the provider returns them; avoid a catch-all rule that retries every exception.
Put a deadline around the work, including the wait
An attempt timeout and a URL budget answer different questions. The first limits how long a single request may occupy a worker. The second limits the total investment in that source, including previous attempts, reading response bodies, and backoff. Add a separate batch deadline when the caller needs a timely partial answer.
Use a monotonic clock for elapsed budgets. Before starting an attempt, cap its timeout at the remaining URL budget. Before sleeping, check whether the delay leaves time for another useful attempt. If the server asks you to wait longer than the remaining budget, record a deferred or exhausted result instead of shortening its delay and immediately retrying.
Retry-After can contain seconds or an HTTP date. HTTP semantics and retry rules also distinguish operations that may be repeated safely from requests with uncertain effects. A timed-out request may already have reached the service. Do not infer that retrying a billable API request is free, or that repeating an arbitrary POST cannot duplicate work.
Backoff spreads attempts over time; jitter varies their timing across workers. Set both an attempt cap and a delay cap. AWS's retry-with-backoff guidance describes transient-error handling and the added load and latency retries can create. Choose your actual settings from measured latency, provider limits, the caller's deadline, and the value of another attempt. A sample constant is not a production recommendation.
Run a small example that keeps partial results
The example below uses the public GET interface documented in AnyCrawler's free URL-to-Markdown tool. It requires no API key and is suitable for checking response handling. Free results can be cached, so this exercise does not demonstrate a fresh fetch on every call or production API billing behavior.
Save the code as example.mjs and run it with a Node.js runtime that provides fetch and AbortSignal.timeout. The demonstration makes two reads of the same permitted example page: one checks for its actual heading, while the other deliberately requires a nonexistent section. This isolates the content-validation failure without intentionally overloading a service or provoking access restrictions.
The policy allows at most three attempts, an attempt timeout of 15 seconds, and a URL budget of 45 seconds. The jitter window starts at 500 milliseconds and is capped at 4 seconds. These are explicit demonstration settings exercised by the local fixtures, not measured optimal values. Only failures identified as potentially transient by this example's transport policy or by an explicit response flag receive another attempt.
// Example policy values, not universal service recommendations.
const policy = { attempts: 3, timeoutMs: 15000, budgetMs: 45000,
baseMs: 500, capMs: 4000 };
const transient = new Set([408, 429, 500, 502, 503, 504]);
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));
export function retryAfterMs(value, now = Date.now()) {
if (!value) return 0;
if (/^\d+$/.test(value.trim())) return Number(value) * 1000;
const date = Date.parse(value);
return Number.isFinite(date) ? Math.max(0, date - now) : 0;
}
export async function readOne(url, expected, options = {}) {
const p = { ...policy, ...options.policy };
const request = options.request ?? fetch;
const pause = options.pause ?? sleep;
const now = options.now ?? (() => performance.now());
const random = options.random ?? Math.random;
const end = now() + p.budgetMs;
const history = [];
const finish = (state, extra = {}) => ({ url, state, history, ...extra });
for (let attempt = 1; attempt <= p.attempts; attempt++) {
const remaining = end - now();
if (remaining <= 0) return finish("budget_exhausted");
let wait = 0;
const event = { attempt, observedAt: new Date().toISOString() };
history.push(event);
try {
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", url);
const response = await request(endpoint, {
signal: AbortSignal.timeout(Math.max(1,
Math.floor(Math.min(p.timeoutMs, remaining))))
});
event.httpStatus = response.status;
event.requestId = response.headers.get("x-request-id");
wait = retryAfterMs(response.headers.get("retry-after"));
const raw = await response.text(); // deadline also covers body reading
if (!response.ok) {
event.reason = "gateway_http_error";
if (!transient.has(response.status)) return finish("permanent");
} else {
let data;
try { data = JSON.parse(raw); }
catch { event.reason = "invalid_json"; return finish("content_invalid"); }
if (!data || typeof data !== "object") {
event.reason = "invalid_shape"; return finish("content_invalid");
}
event.targetStatus = data.status_code ?? null;
event.finalUrl = data.final_url ?? null;
event.creditsUsed = data.credits_used ?? null;
event.reason = data.error_code ?? null;
const targetOk = Number.isInteger(data.status_code)
&& data.status_code >= 200 && data.status_code < 300;
if (data.ok === false || !targetOk) {
if (data.retryable !== true) return finish("needs_review");
} else {
const markdown = data.results?.markdown;
if (typeof markdown !== "string" || !markdown.includes(expected)) {
event.reason = "expected_content_missing";
return finish("content_invalid");
}
return finish("success", { finalUrl: data.final_url ?? null,
title: data.results?.title ?? null, markdown,
responseTimestamp: data.timestamp ?? null });
}
}
} catch (error) {
event.reason = error.name;
const code = error.cause?.code;
const network = ["ECONNRESET", "ETIMEDOUT", "EAI_AGAIN",
"UND_ERR_CONNECT_TIMEOUT", "UND_ERR_SOCKET"].includes(code);
if (!["TimeoutError", "AbortError"].includes(error.name) && !network)
return finish("needs_review");
}
if (attempt === p.attempts) return finish("retry_exhausted");
const jitter = random() * Math.min(p.capMs, p.baseMs * 2 ** (attempt - 1));
const delay = Math.max(wait, jitter);
if (delay >= end - now()) return finish("budget_exhausted");
event.delayMs = Math.ceil(delay);
await pause(Math.ceil(delay));
}
}
export async function demo() {
const jobs = [
["https://example.com/", "Example Domain"],
["https://example.com/", "THIS_REQUIRED_SECTION_IS_ABSENT"]
];
const results = [];
for (const [url, expected] of jobs) results.push(await readOne(url, expected));
return { complete: results.every(row => row.state === "success"), results };
}
console.log(JSON.stringify(await demo(), null, 2));
In the live run used to prepare this article, both requests returned outer HTTP 200 and target status 200. The first produced the title Example Domain and the expected Markdown. The second stopped as content_invalid because its required phrase was absent. Each used one attempt. The response did not supply final_url or credits_used; both were retained as null rather than invented. This validates the partial-result path on that example page, not a service-wide success rate.
The important part of the resulting application record is:
{
"complete": false,
"results": [
{ "state": "success", "title": "Example Domain" },
{ "state": "content_invalid", "reason": "expected_content_missing" }
]
}
This shortened view omits Markdown and attempt history for readability. The executable version retains them. Keep the successful document available to downstream processing; describe which requested evidence is still missing before producing any answer that depends on it.
Test the failure branches without manufacturing an outage
A live success cannot prove that retry code stops correctly. Use controlled responses and a fake clock to exercise branches that would be disruptive or unreliable to trigger against a public service.
The local checks for this example covered a temporary server failure followed by success, a throttling response with a delay header, an unchanged permission failure, repeated gateway timeouts, an attempt timeout followed by success, and a requested wait longer than the budget. Additional cases checked target errors inside successful gateway responses, missing content, explicit retryability, malformed JSON, and an unrelated programming exception. All eleven scenarios passed their expected state and attempt-count assertions. The date and seconds forms of Retry-After were also checked. Those results describe the fixtures; no live outage or throttling benchmark was performed.
Before using the pattern in a larger pipeline, verify these invariants:
- A terminal outcome exists for every submitted input, including items never started before the batch deadline.
- A failed input cannot erase or overwrite an already accepted document.
- The worker does not sleep past a deadline and then start another request anyway.
- A cancellation reaches in-flight requests, while durable attempt records survive the cancellation.
- Retrying an unresolved item retains its task identity, prior evidence, and an explicit refresh policy.
The example runs sequentially to keep its scope clear. A production queue needs bounded concurrency, limits shared across workers, cancellation, a response-size limit, and durable checkpoints. If an SDK already retries, inspect that behavior before adding another retry layer. Independent caps can otherwise multiply the work and hide the real attempt count.
Carry the same contract into Fetch and Render
For controlled known-URL retrieval, use the request and response fields in the AnyCrawler Fetch API guide. It provides the production integration path, including explicit mode selection and usage metadata. When adapting the example, map outer status, target status, provider error details, request ID, and reported credits into separate fields. The free demo is not a drop-in production POST adapter.
A missing section should trigger diagnosis before escalation. Verify that the submitted URL is the intended document and that the extraction rule matches the task. If the text only appears after JavaScript runs, Render is a different retrieval strategy to evaluate under its own budget. Repeating Fetch cannot create content that is absent from the returned HTML. Screenshot can retain visible context for review; it does not establish that all required text was extracted.
Search discovers candidate URLs. Fetch retrieves an HTML document. Render reads a page after browser execution. Full browser automation is a separate requirement when a permitted workflow needs interaction. Retrying or changing retrieval mode does not grant access to restricted material, and this architecture does not imply that AnyCrawler provides batch scheduling, change detection, or alerts.
Which missing source should hold up the answer?
The difficult stopping decision is often about evidence value. A failed optional background page may leave a useful answer possible; a missing primary specification can leave the central claim unsupported. Track source outcomes alongside the claims they were meant to establish, so a large number of successful pages cannot conceal a critical gap.
Before increasing retries, choose one real task and label its required evidence. Run the small validator on a permitted URL, inspect the failure record, then decide whether the next attempt should wait, change retrieval strategy, or stop for review. The remaining design question is how your application distinguishes an acceptable partial answer from a result that should be withheld until a specific source is verified.
Frequently asked questions
How many times should a web scraper retry?
There is no single count that fits every workload. Set a maximum attempt count together with a total URL budget, and account for limits enforced by the provider. The example uses three attempts only to demonstrate a bounded loop. A fast interactive request and a deferred research job may justify different policies. Review actual failure classifications and the value of the missing source before raising the limit.
Should every HTTP 500 or 503 response be retried?
Treat a server status as a starting signal, then inspect the documented error meaning. Temporary overload can justify another attempt within budget, while a persistent configuration problem usually needs intervention. A provider's explicit non-retryable failure should not be overridden by a generic status rule. Preserve the final error and request identifier so repeated failures can be diagnosed without continuously resubmitting the same work.
Why can a request return 200 and still fail the task?
The outer response can successfully deliver an API result even when the target page failed. A retrieved document can also lack the content your task requires. Validate the target status and the required document fields separately. In the demonstration, the second request reached the example page successfully but intentionally required an absent section, so the application kept a content-invalid result instead of accepting unusable evidence.
What should happen when Retry-After exceeds the budget?
End the current item with a recorded reason instead of retrying before the requested delay. Your application can place it in a deferred queue if a later attempt still has value and scheduling is available. Preserve the requested delay and prior attempts for that decision. The demonstration marks the item budget-exhausted; it does not implement a scheduler or promise that a later request will succeed.
Does timing out a request mean it used no credits?
No. A client timeout tells you that the client stopped waiting; it does not establish what the remote service completed or charged. Keep any returned usage fields and request identifiers, then reconcile uncertain usage using the provider's available records. The free demonstration returned no credits field, so it stored null. Do not interpret that missing field as a measured zero or a production billing guarantee.
How do I retry failed URLs without losing successful pages?
Store each input's state independently and retain accepted content with its provenance. Build a later queue from the unresolved states that still justify another attempt, carrying forward task identity and failure history. Keep content-validation problems separate from transient transport errors, since they often require different action. Before generating an answer, check whether the remaining failures affect required claims rather than assuming that any partial batch is sufficient.






