Before an AI agent fetches a public page, check the robots.txt file for that page's scheme, host, and port using the crawler identity that will make the request. Apply the matching user-agent group and its most specific Allow or Disallow rule to the target path. If the path is disallowed, skip it; if the rules cannot be reached, defer the read. Record the target, identity, rule outcome, and check time with the result. This preflight addresses crawler access rules. It does not replace authentication, a site's other access conditions, or a second check if the request moves to another origin.
Match the rule to the request you will actually send
A robots decision is about a particular request, not a domain name in the abstract. https://example.com/robots.txt governs that HTTPS origin; a different subdomain, scheme, or nonstandard port needs its own file. The file belongs at the root, not beside the page. The Robots Exclusion Protocol defines how a crawler selects a user-agent group and compares its Allow and Disallow paths. If more than one group matches the crawler's product token, their rules are combined; otherwise the * group applies when present.
Use the identity of the client making the outbound request. Your application's name and the HTTP User-Agent sent by a delegated fetch service may differ. A check for MyResearchBot cannot silently stand in for a request made under another identity. Choosing a different crawler's name because its group is more permissive would misstate the request. If you cannot establish which identity will fetch the page, mark the preflight unresolved rather than selecting a favorable group.
For example, Disallow: /reports/ and Allow: /reports/public/ can coexist. A request for /reports/public/index.html takes the longer Allow match. A request for /reports/private.html remains disallowed. An empty or unmatched rule set has different meaning from a failed attempt to retrieve the rules file. Treating both as “no restriction found” hides the failure that an agent needs to report.
Give every preflight a state
The decision should be explicit enough for your worker to stop before it calls a page-reading API. This is a conservative application policy built on the protocol, with separate states for a rule decision and a rules-file failure:
/robots.txt observation |
Decision for the target URL | Record and next action |
|---|---|---|
| Successful read; applicable rule allows the path, or no rule matches | allow |
Save origin, crawler identity, target path, rules retrieval time, and matching rule; continue to other access checks. |
| Successful read; the most specific applicable rule disallows the path | deny |
Save the matched rule and skip Fetch or Render for this URL. |
404 for the rules file |
allow-by-policy only if your application's policy accepts missing rules |
Save the 404 separately from an explicit Allow; check other access constraints before reading. |
401, 403, or 429 for the rules file |
defer under this conservative policy |
Do not turn an authentication or rate signal into automatic permission; investigate or retry when appropriate. |
5xx, timeout, DNS failure, or an unreadable response |
defer |
Keep the failure and retry later within a bounded policy; do not guess a path rule. |
The RFC says a crawler may access resources when the rules file is unavailable with a 4xx response. It also says a rules file unreachable through server or network errors means complete disallow while that condition holds, subject to its long-outage and cache provisions. Choosing to defer on 401, 403, and 429 is the workflow's stricter safety decision, not a claim that the RFC treats every 4xx as a ban. Google's published interpretation makes different crawler-specific choices, including special handling for 429; do not copy those choices into an agent without deciding what your own client should do.
Fetch the rules file with a normal GET and evaluate the target's path under the original origin. If that rules request redirects, follow a bounded redirect policy and apply the resulting rules to the original authority, as the RFC specifies. Store whether the rules came from a redirect so a later audit can reconstruct the decision. An Allow result is a prerequisite for this workflow, not an instruction to send unlimited requests.
Read an allowed page, then inspect what came back
The public https://anycrawler.com/robots.txt currently has a wildcard Allow for /. That makes an AnyCrawler Blog page a suitable, owned example for a one-page probe. After checking the rule, the following command calls the no-key endpoint documented on the Free URL to Markdown tool:
curl --silent --show-error --get "https://api.anycrawler.com/free/v1/crawl" --data-urlencode "url=https://anycrawler.com/blog/chunk-web-pages-for-rag-with-citations/"
In the recorded run, the request returned HTTP 200. The JSON had ok: true, a target status_code of 200, the submitted requested_url, a page title, and nonempty Markdown. It did not include final_url, canonical_url, or credits_used. Store those fields as unknown when absent; do not fill them from the submitted URL or from a sample response in documentation. The free route may return cached content, so this run demonstrates a readable public page, not a fresh crawl guarantee or proof of the service's outbound user agent.
For a production integration with an API key, the known-URL Fetch API documents POST /v1/crawl/page with method=fetch and the response fields to inspect. Keep your allow/deny/defer record before that call, then save the response status and any URL identity or usage fields actually returned. Search discovers candidate URLs; Fetch reads one selected URL; Render is for content that requires browser execution; Screenshot records visible state when that matters. None of these output choices answers the policy question for you.
A compact evidence record can hold requested_url, checked_origin, crawler_identity, robots_status, matched_rule, decision, checked_at, fetch_status, and observed_final_url. A null final URL means “not observed,” not “same as requested.” Keep the robots observation and page-read result separate so a later success response cannot erase the earlier decision.
A redirect can change the answer
A redirect while retrieving /robots.txt is different from a redirect of the target page. The former supplies rules for the original authority when handled according to the protocol. The latter may move the page read to a new scheme, host, or port whose rules have not been checked. A URL shortener or an old documentation address can create that change even when the original path was allowed.
If your fetch layer exposes redirect hops, pause before following a hop to a new origin, obtain that origin's robots file, and evaluate the new path for the actual outbound identity. If the layer follows redirects internally and returns only the eventual page, a preflight on the submitted URL alone cannot prove that every visited origin was allowed. Require redirect controls or a trustworthy redirect trace when that proof is part of your task; otherwise retain the uncertainty and do not label the entire path policy-checked.
What an Allow rule leaves open
The protocol communicates crawler access preferences. It does not authenticate a client, protect content, settle a site's terms, or establish permission to use the content for a particular purpose. The RFC's security discussion explicitly warns that listing a path in robots.txt is not access control. If a page requires login, a paywall, or another authorization boundary, an Allow rule does not remove it. Likewise, a page that loads in a browser is not automatically approved for automated collection.
Operational limits still matter after an Allow. A target can return a rate signal or an error, and a successful HTTP response can contain an empty or wrong document. Handle those as separate states. Limit requests, validate the expected content, and preserve the failure instead of prompting an agent to keep retrying until it finds a route around the restriction.
Who can attest to the final request?
The hardest boundary appears when the application asks another service to fetch the page. The application knows its submitted URL and its own intended policy, but the service controls the outbound identity and may follow redirects. A preflight can be exact only if those details are known or exposed at the point of the request. If your product cannot reveal them, what evidence is sufficient to approve the read: a restricted allowlist of destinations, an audited fetch gateway, or a deliberate defer state? That choice belongs in the application's policy and evidence contract before an agent starts relying on returned text.
Frequently asked questions
Does a public page mean a crawler may fetch it?
No. Public visibility says that someone can load a page without a private session; it does not tell you which automated clients the site asks to keep away. Check the /robots.txt file for the page's origin and the identity that will make the request. Then apply the relevant path rule. Authentication, site terms, rate pressure, and the intended use of the content are separate questions. A page's HTTP 200 response cannot answer those questions for your agent.
What if a site has no robots.txt file?
An actual 404 for /robots.txt is a different observation from a timeout or a server error. RFC 9309 permits a crawler to access resources when the rules file is unavailable with a 4xx response, while leaving room for application policy. If your workflow allows a read after a 404, record it as allow-by-policy, not as an explicit Allow rule. The other access checks still apply, and a different origin needs its own check.
Should the agent retry when robots.txt returns an error?
It depends on the failure. A server error or network failure leaves the rules file unreachable; the protocol calls for treating that condition as disallow, with provisions for persistent outages and cached rules. This workflow defers the page read and retries the rules request within a bounded policy. It also defers on 401, 403, and 429 rather than interpreting them as a green light. Keep the response code, time, and retry result so the agent cannot turn a transient failure into a silent allow.
Which user-agent group should I evaluate?
Evaluate the group for the client that actually sends the page request. The protocol matches a crawler product token, combines groups matching that token, and falls back to the wildcard group when there is no matching named group. If an external Fetch service sends the request, your application label may not be that service's outbound identity. Verify what it uses or retain an unresolved state. Do not select another crawler's group because its paths happen to permit the URL.
Is one robots check enough if the page redirects?
Only while the request remains within the origin whose rules you checked. A redirect of the robots file can still yield rules for the original authority under the protocol. A redirect of the page to a different scheme, host, or port creates a new target that needs its own policy decision. When a service hides the redirect chain, the submitted URL's preflight cannot establish what happened at the final origin. Require visibility or record that limit before claiming the read was policy-checked.
Does a successful AnyCrawler result prove robots compliance?
No. The tested no-key request confirms that one owned, publicly allowed page returned Markdown and status fields. It does not reveal every outbound hop or independently establish the fetch service's user agent; this run also lacked a final URL field. Use the result to verify extraction, then keep the robots preflight as its own record. For a strict workflow, ensure the component that knows the actual outbound identity and redirect path can enforce or attest to the decision.






