How to Extract Text from a Website for AI and Automation
To extract text from a website for AI or automation, start with a known public URL, fetch the page, and verify that the returned body contains the main content, not just navigation, scripts, or a JavaScript shell. Save the requested URL, final or canonical URL when available, HTTP status, title, Markdown, and extraction time with the result. If the useful text is missing because the page builds content in JavaScript, switch from a lightweight fetch path to a browser render path. For a quick one-off check, try the free crawler; for production control, compare Fetch and Render.
Text extraction sounds simple until the output reaches an agent. A human can ignore a menu, cookie banner, or blank app shell. A model usually cannot. The practical goal is not "all text on the page"; it is a reproducible page record that keeps useful content, preserves enough structure for downstream use, and tells the next step whether the crawl succeeded.
What "website text" means in a crawl result
There are at least four different outputs that people call website text:
| Output | What it captures | Good for | Common failure |
|---|---|---|---|
| Visible text | Text a browser user could select or read | Quick inspection, UI review | Can include menus, repeated chrome, banners, and layout-only labels |
| Main content | The central article, doc, product copy, or page body | Research, RAG, summarization, extraction | May miss important sidebars, tables, disclosures, or related links |
| Markdown | Main content plus headings, lists, links, and code in a compact format | AI context, storage, diffing, source review | Bad conversion can flatten tables, code blocks, or nested structure |
| Metadata and links | Title, description, canonical hints, outbound links, media, and provenance fields | Citation, deduplication, follow-up crawling | Not a substitute for the body text |
Browser APIs show why raw text is not a complete strategy. MDN describes innerText as rendered text, while textContent ignores rendered appearance and may include hidden content or scripts in ways a user would not see (MDN innerText). That distinction is useful for debugging, but an AI ingestion workflow usually needs a higher-level extraction contract: which page was requested, what content was considered primary, and how success was verified.
A reliable extraction workflow
Use a small, auditable sequence before you send page text into a model:
- Start with one public URL. Do not treat a single-page fetch as a whole-site crawl.
- Fetch the page and keep the HTTP status with the output.
- Inspect
results.titleandresults.markdownbefore assuming the extraction is useful. - Check for expected phrases, headings, tables, code blocks, or links that matter to the task.
- Keep requested URL, final URL or canonical URL when returned, timestamp, and crawl metadata with the text.
- If the main body is missing, test whether the page needs JavaScript rendering before retrying blindly.
The most common mistake is passing any non-empty string to the next agent. Boilerplate can be long enough to look successful. A short but complete documentation page can look suspiciously small. The validation should come from the task: a product comparison needs plan names or price context; a documentation extraction needs headings and code; a research source needs enough body text to support the citation.
Tested free-endpoint example
The following example was run on August 25, 2026 against AnyCrawler's public free endpoint. It uses three permitted public pages: one MDN article page, one Google documentation page, and one React documentation/product-style page. The free endpoint returned HTTP 200 and Markdown for all three inputs during this run. It did not expose a credits field, so the test ledger records credits as not_exposed instead of inventing a value.
const inputs = [
"https://developer.mozilla.org/en-US/docs/Web/API/HTMLElement/innerText",
"https://developers.google.com/search/docs/crawling-indexing/javascript/javascript-seo-basics",
"https://react.dev/learn",
];
for (const url of inputs) {
const endpoint =
"https://api.anycrawler.com/free/v1/crawl?url=" + encodeURIComponent(url);
const response = await fetch(endpoint, {
headers: { accept: "application/json" },
});
const data = await response.json();
console.log({
input: url,
status: data.status_code,
title: data.results?.title,
markdownLength: data.results?.markdown?.length ?? 0,
timestamp: data.timestamp,
credits: data.credits_used ?? "not_exposed",
});
}
| Test input | Result on August 25, 2026 | What was checked |
|---|---|---|
MDN innerText article |
status_code: 200, title returned, Markdown length 30,748 characters |
Article-style page returned a readable body |
| Google JavaScript SEO documentation | status_code: 200, title returned, Markdown length 49,275 characters |
Documentation page returned headings and body content |
| React Learn page | status_code: 200, title returned, Markdown length 27,679 characters |
JavaScript-heavy documentation/product page still exposed extractable content through the free path |
These results prove only those requests at that time. They do not prove that every page on those domains will extract cleanly, that the free endpoint is uncached, or that a production integration should ignore the authenticated crawl API. The free page is best for a quick inspection. Production workflows should use explicit fetch or render mode, store request IDs and usage fields when available, and define failure handling before the agent depends on the content.
When to use Fetch and when to Render
Use Fetch when the main content is already present in the returned HTML. That covers many public articles, documentation pages, reference pages, help centers, and simple landing pages. It is the cleaner default when you already have a URL and need text quickly.
Use Render when the first response is mostly an app shell, the route changes in the browser, or the target content appears only after JavaScript runs. Google's JavaScript crawling guidance makes the same operational distinction: classical server-rendered pages can be read from the HTML response, while app-shell pages require JavaScript execution before generated content is visible to the renderer (Google Search Central JavaScript SEO basics).
| Signal | Start with Fetch | Upgrade to Render |
|---|---|---|
| Page type | Blog, docs, article, changelog, reference page | Single-page app, client route, browser-built table |
| Markdown body | Contains expected headings and main text | Empty, generic, or mostly navigation |
| Required evidence | Text and links are enough | Visible browser state changes what the agent can trust |
| Failure handling | Retry only finite network or upstream failures | Render once with clear wait criteria, then validate again |
For AnyCrawler, the next decision is explicit: use the fetch page API for a known URL whose HTML already contains the useful body, and the render page API when JavaScript execution is part of the content contract.
The integrity checklist before an agent uses the text
Before a model summarizes, cites, embeds, or extracts fields from page text, validate the crawl result against a small checklist:
- Access: HTTP status is success, and the result is not an error page, login page, or blocked placeholder.
- Identity: requested URL, final URL, canonical URL when returned, and title are stored together.
- Completeness: expected headings, facts, sections, or entities appear in Markdown.
- Structure: lists, tables, code blocks, and links that matter to the task did not flatten into unreadable text.
- Noise: repeated navigation, footer links, cookie text, unrelated recommendations, and ads are not dominating the output.
- Freshness: the capture timestamp and any cache behavior are visible enough for the downstream decision.
- Boundary: the result is labeled as one page, not a site crawl, monitor, alert, or source-of-truth database.
This is also where a screenshot can help. Markdown records the textual body; a screenshot records visible state. If the task depends on layout, disclosure placement, a pricing card, or a page state that a reviewer must see, pair text extraction with screenshot evidence instead of asking the Markdown to prove everything.
A source record schema that keeps automation honest
For automation, store text as part of a source record rather than a bare string:
{
"requestedUrl": "https://example.com/docs",
"finalUrl": "https://example.com/docs",
"canonicalUrl": "https://example.com/docs",
"statusCode": 200,
"capturedAt": "2026-08-25T02:07:17.205Z",
"crawlMode": "fetch",
"title": "Example Docs",
"markdown": "# Example Docs\n\n...",
"validation": {
"expectedPhrasesPresent": true,
"mainContentLooksComplete": true,
"needsRenderRetry": false
}
}
That schema does not need to be large. It needs to make failure visible. If the extraction fails, keep the failed record with statusCode, error fields, retryability, and the reason it was rejected. A research or monitoring agent can then continue with other sources without silently mixing valid evidence and broken fetches.
The open boundary: extraction quality is task-specific
There is no universal threshold for "enough text." A 600-word policy page can be complete, while a 5,000-word Markdown output can be mostly navigation and related links. The unresolved decision is how strict each workflow should be. A RAG ingestion pipeline may need chunk-level structure, link preservation, and stable canonical identity. A support triage agent may only need the title, status, and a short body section. A compliance review may need both Markdown and screenshot evidence.
Treat text extraction as a contract with the downstream task. Define what must be present, what can be omitted, and when the agent should stop rather than reason from a weak page record.
Frequently asked questions
What is the fastest way to extract text from a website?
The fastest practical path is to start with a single public URL and run a lightweight fetch. If the returned Markdown includes the expected title, headings, and main body, you can use it without browser rendering. Speed should not be the only check, though. A fast response that contains navigation, an error page, or an empty app shell is not useful evidence. Validate content completeness before sending the text to an AI workflow.
Is textContent enough for website text extraction?
textContent is useful when you control the DOM fragment and only need raw node text, but it is too low-level for most web extraction workflows. It does not decide what the main article or document body is, and it can include hidden or irrelevant text depending on the source. For AI use, you usually need cleaned Markdown, source identity, status fields, and validation checks rather than a single DOM property.
When should I render a page instead of fetching it?
Render the page when the useful content is created after JavaScript runs, appears through client-side routing, or depends on browser state. Start with fetch when the page is an ordinary article, documentation page, changelog, or other HTML-first source. The decision should be based on the output, not on the framework. If fetch returns the expected main body, rendering may only add cost and complexity.
Can extracted website text be used directly for RAG?
It can be a good input, but extraction alone does not solve retrieval quality. You still need source records, canonical or final URL handling, chunking, deduplication, metadata, and validation that the relevant content was captured. Markdown is often easier for RAG than raw HTML because headings and lists survive in a compact form. It is still only the transport format, not a guarantee that retrieval will answer correctly.
How do I know whether the extracted text is complete?
Check the result against the task. Look for required headings, expected entities, tables, code blocks, links, or phrases that the downstream workflow needs. Also verify that the title and status code match the target page, and that the Markdown is not dominated by menus or footer links. If the expected content is missing, try render mode once with a clear wait condition, then validate the rendered result again.
Does AnyCrawler extract text from a whole website?
This article covers extracting text from a single public page URL. AnyCrawler's current public pages distinguish search, page fetch, render, and screenshot workflows, and the free endpoint is a quick single-URL test. Do not treat one page extraction as a recursive site crawl, scheduler, diff engine, alerting system, login automation tool, or paywall bypass. If the URL is unknown, start with search, then fetch or render selected pages.






