How to Scrape Dynamic Web Pages: When to Fetch and When to Render
Start with Fetch, then upgrade to Render only when a field you actually need is missing from the returned HTML. The presence of JavaScript, a React bundle, or an app-shaped URL is not enough to justify a browser. Define expected content first, fetch the page, and validate its title, status, final identity, and required markers. Render when client-side routing, asynchronous data, or post-load DOM changes prevent that contract from passing. A successful request is not the finish line; the extracted result must contain the evidence your workflow expects.
Dynamic does not always mean browser-only
Client-side rendering can generate page content in the browser, but modern sites also mix server output with client behavior. A page may ship complete article text in its first HTML response and use JavaScript only for menus, analytics, or interaction. That page is dynamic in implementation yet still suitable for a lightweight Fetch.
Hydration makes this distinction especially important. React's official hydrateRoot documentation explains that components can attach to HTML already generated on the server. In that case, the initial response may contain everything an extraction workflow needs even though the browser later activates the interface.
The practical question is therefore not "Does this page use JavaScript?" It is "Does the response already contain the specific content and identity fields required by this task?"
| Observed result after Fetch | What it means | Next action |
|---|---|---|
| Required title, body, links, and identifiers are present | JavaScript is not blocking this extraction contract | Keep Fetch |
| Only a shell, loading message, or navigation frame is present | Useful content may be created in the browser | Try Render |
| Main content exists, but one optional widget is absent | The page is partially client-rendered | Render only if that widget is required |
| Requested URL resolves to another document | Redirect or route identity needs review | Save requested, final, and canonical URLs before deciding |
| Status is 200 but required markers are missing | Transport succeeded; extraction did not | Treat as incomplete, not successful |
Prove the gap with a Fetch-first test
On August 29, 2026, we ran the public, no-key AnyCrawler Fetch endpoint against two permitted pages with the same request shape. The React hydrateRoot documentation returned HTTP 200, the expected title, 36,723 Markdown characters, and the marker hydrateRoot. The public AJAX scraping practice page also returned HTTP 200, but its 1,246-character Markdown did not contain Spotlight. In a browser, selecting 2015 added a table containing Spotlight to that same page.
The free response reported the requested URL and target status but did not include final_url or credits_used; those fields must be recorded as unavailable rather than invented. The browser check proves that the target content appears after interaction. It does not claim that an authenticated AnyCrawler Render request was run in this test.
This is the exact Node.js probe used for the Fetch half of the comparison:
const probes = [
{
label: "hydrated documentation",
target: "https://react.dev/reference/react-dom/client/hydrateRoot",
marker: "hydrateRoot",
},
{
label: "AJAX practice page",
target: "https://www.scrapethissite.com/pages/ajax-javascript/",
marker: "Spotlight",
},
];
for (const probe of probes) {
const endpoint = new URL("https://api.anycrawler.com/free/v1/crawl");
endpoint.searchParams.set("url", probe.target);
const response = await fetch(endpoint);
const data = await response.json();
const markdown = data.results?.markdown ?? "";
console.log({
label: probe.label,
gatewayStatus: response.status,
targetStatus: data.status_code ?? null,
requestedUrl: data.requested_url ?? null,
finalUrl: data.final_url ?? null,
title: data.results?.title ?? null,
markdownChars: markdown.length,
markerPresent: markdown.includes(probe.marker),
creditsUsed: data.credits_used ?? null,
});
}
Run the probe with current public URLs and choose markers that represent the real task: a product name, table heading, article section, price label, or record identifier. A generic word such as Home is too weak because navigation can satisfy it while the main content is still absent.
Use an extraction contract, not a page-type guess
Before calling either path, write down what a usable result must contain. This turns "the page looked dynamic" into a reproducible decision.
For a documentation page, the contract might require the H1, a named API method, at least one code block, the final URL, and a successful target status. For a pricing page, it might require plan names, billing periods, currency, and the disclosure text that qualifies the price. For a client-side route, it should include the route-specific heading rather than a shared application-shell title.
A compact validation record can use these fields:
{
"requested_url": "https://example.test/app/pricing",
"final_url": null,
"canonical_url": null,
"method": "fetch",
"target_status": 200,
"required_markers": ["Plans", "Billing period", "Terms"],
"missing_markers": ["Billing period", "Terms"],
"content_state": "incomplete",
"next_action": "render"
}
Do not silently turn missing fields into empty strings and continue. Preserve null, the method used, and the failed markers so another worker—or a human reviewer—can see why the pipeline escalated.
Recognize the failure patterns that justify Render
Hydration can still be Fetch-friendly
A hydrated page may send a complete HTML snapshot, then attach interaction logic. If the Fetch result passes the contract, rendering adds browser cost without adding useful evidence. The React documentation test above is a concrete example: the page uses hydration, yet the required documentation content was already extractable.
Client-side routes can return a shared shell
Some applications serve the same root document for many URLs and let JavaScript resolve the route. A Fetch can therefore return HTTP 200 and a valid-looking title while omitting the route-specific record. Validate a heading or identifier unique to the requested route. If that marker appears only after browser navigation, use Render for that URL pattern.
Asynchronous sections may appear after the document is ready
The browser's DOMContentLoaded event means the HTML has been parsed and deferred scripts have executed; it does not wait for every asynchronous script or later data request. A page can pass a load milestone while the required table is still absent. The extraction contract, not a generic load event, should decide readiness.
Redirects can hide an identity error
Save the requested URL, final URL, and canonical URL separately when the API returns them. A redirect to a sign-in page, regional version, consent screen, or generic app shell may still produce HTTP 200. Do not cite or index the content until its final identity matches the intended source. If the current response omits a field, record that absence explicitly.
Timeouts need a bounded failure state
A timeout does not tell you whether the target was permanently unavailable, still loading, or partly extracted. Store the method, wait condition, elapsed boundary, visible markers, and retryability signal. Retry only a finite number of transport or temporary failures. Repeating the same browser request without changing the condition usually spends more time without improving the evidence.
Render with a content-specific readiness rule
When Fetch fails the contract, switch the same known URL to the JavaScript Render API and wait for evidence that matters to the task. Prefer a route-specific heading, a stable result row, or a required data label over a fixed sleep.
This aligns with Playwright's own guidance: its load-state documentation discourages treating networkidle as a universal test of readiness and recommends assertions against the page instead. A polling dashboard or analytics stream may never become truly idle, while a required table can be ready much earlier.
Use this Render checklist:
- Keep the original requested URL and the final browser URL.
- State the required selectors, phrases, or record identifiers before the request.
- Bound the wait and return a structured incomplete state when the markers never appear.
- Record target status, extracted title, Markdown, and any cache or usage fields the response actually supplies.
- Verify redirects and canonical identity before accepting the document.
- Keep screenshots separate when visible layout is part of the evidence; rendered text alone does not preserve appearance.
If the browser must log in, click through a workflow, submit a form, or maintain session state, the task has moved beyond read-only rendering into browser automation. Render is an extraction path, not a general interaction agent.
Put Fetch and Render behind one bounded state machine
A production workflow should make escalation explicit:
- FETCH_PENDING: request the known public URL through Fetch.
- FETCH_VALIDATE: compare title, status, identity, and required content markers with the extraction contract.
- FETCH_ACCEPTED: store the result when all required markers pass.
- RENDER_PENDING: escalate only when missing content is plausibly browser-created.
- RENDER_VALIDATE: apply the same contract to the rendered result, plus final browser URL and wait evidence.
- RENDER_ACCEPTED: store the result with method and provenance.
- INCOMPLETE: stop with structured missing markers, error class, and retryability after the bounded attempt fails.
This state machine prevents two common mistakes: accepting a 200 shell as complete and sending every page to a browser by default. It also makes partial failure auditable because the system retains the rejected Fetch result and the reason for escalation.
Start with the Fetch Web Page API for known public URLs, or use the free URL-to-Markdown crawl to inspect one page without an API key. If the URL itself is unknown, the next decision is discovery rather than rendering; use the broader Search, Fetch, Render, or Screenshot guide to choose that path.
What should trigger Render when pages keep changing?
The difficult boundary is not technical branding such as "SPA" or "React app." It is the stability of the extraction contract. A page can move from client-only rendering to server output without changing its URL, or add one asynchronous widget while leaving its main document Fetch-friendly. Workflows should therefore observe which required markers fail by route pattern and revisit the rule when the site changes.
The open question is how specific those markers should be. Broad markers are stable but may accept incomplete pages; narrow selectors catch omissions but break when presentation changes. The useful compromise is semantic: validate durable names, record identifiers, and content sections, then review the failure history before promoting any route to Render-by-default.
Frequently asked questions
Can I detect dynamic content by looking for script tags?
No. Script tags show that a page loads JavaScript, not that JavaScript creates the content your workflow needs. Hydrated and server-rendered sites may ship a complete document and use scripts only to add interaction. Define required markers, run Fetch, and inspect the returned Markdown and identity fields. Escalate only when those markers are absent for a reason browser execution can plausibly resolve.
Should I always try Fetch before Render?
For a known public URL, Fetch-first is a strong default because it gives you a cheap diagnostic result and a reasoned escalation path. It is not mandatory when prior, current evidence proves that a route is client-only and your required fields never exist in initial HTML. Even then, keep the contract and periodically recheck it; a site can change its rendering architecture without changing its URLs.
What does an HTTP 200 response prove?
It proves that the request reached a target that returned a successful HTTP status. It does not prove that you received the intended document, that a client-side route resolved, or that the required content loaded. Check the requested and final identities, title, canonical when available, and task-specific markers. A 200 login page, consent screen, app shell, or loading placeholder should remain an incomplete extraction.
Which wait condition should a renderer use?
Use the earliest condition that reliably exposes the required content, then validate that content directly. DOMContentLoaded may be sufficient for server output, while an asynchronous table may need a visible row or heading. A fixed delay is fragile, and network idle is not a universal readiness signal. Bound the wait, name the expected marker, and return an explicit incomplete state if it never appears.
How should I handle redirects and client-side routes?
Store the requested URL, final URL, and canonical URL as separate fields whenever they are available. Validate a route-specific heading or identifier after navigation, because a shared application shell can return 200 for many paths. If the response lands on authentication, consent, or a generic root route, do not accept it as the requested document. Record the mismatch and choose an authorized next step.
Is Render the same as browser automation?
No. Render opens a page so JavaScript can produce a readable document, then returns extracted content. Browser automation is needed when the task must click, type, log in, preserve a session, submit data, or cause other side effects. Keep those contracts separate. If a page requires interaction merely to reveal public content, document that requirement and use an authorized interaction workflow instead of implying that passive rendering performs it automatically.






