Feature
robots.txt Is a Crawl Policy, Not an AI Browser Block
Separate AI crawlers, user-triggered fetchers and browser agents, then choose controls that actually enforce access to your publishing site.
Impetuous · · 4 Min Read

Not reliably. robots.txt can tell a compliant AI crawler not to fetch your pages. It does not make your server deny access, and it is not a dependable block against an assistant browsing on a person’s behalf. The Robots Exclusion Protocol explicitly says its rules “are not a form of access authorization.” RFC 9309
For publishers, the useful distinction is not “AI versus human.” It is which client requests the content, which policy that client follows, and what your server enforces.
Separate the request types
OpenAI’s documentation distinguishes automatic crawling from user-triggered fetching. It also documents a separate signed cloud-browser client. Treat these as separate policy decisions. Crawler documentation · Cloud-browser documentation
| Request type | What robots.txt can do | Operational implication |
|---|---|---|
| Training crawler, such as GPTBot | Express a training-crawl opt-out to the operator | Useful policy signal; not technical access enforcement |
| Search crawler, such as OAI-SearchBot | Control compliant automatic crawling for search | Decide separately from training; blocking can reduce discovery |
| User-triggered fetcher, such as ChatGPT-User | May not govern the request | OpenAI explicitly says robots.txt rules may not apply |
| Agent-operated browser | Depends on the implementation and operator’s policy | Identify the request path before choosing a block |
OpenAI says OAI-SearchBot and GPTBot settings are independent. Blocking GPTBot does not itself block ChatGPT search crawling. Conversely, sites opted out of OAI-SearchBot will not appear in ChatGPT search answers, although they can still appear as navigational links. OpenAI crawler documentation
A proposed policy for allowing OpenAI search crawling while declining GPTBot training crawling is:
User-agent: GPTBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
This is not a universal AI opt-out. Merge it with your existing rules rather than replacing the file. A named bot group does not inherit the User-agent: * group’s restrictions: preserve any path exclusions you still need within the OAI-SearchBot group. Check the final response served through your CDN, not just the origin file. RFC 9309, group matching
A browser agent is not necessarily a crawler
An assistant operating an ordinary browser can use the same browser session as its user. If navigation happens through that browser’s normal network path, the server may receive ordinary browser requests rather than a distinct crawler identity. A GPTBot rule does not identify that activity.
Tobira’s analysis describes this local-session case as an assistant riding inside the user’s browser, with the user’s cookies, session and network identity. That is a useful architectural distinction—not evidence that every AI browsing product uses that path. Tobira’s analysis
Product names are also unstable policy targets. OpenAI’s Atlas retirement guidance scheduled it to stop working on August 9, 2026, while moving browser-based agentic capabilities into ChatGPT and Codex and directing users toward desktop and Chrome experiences where available. That guidance does not establish how every replacement sends requests. OpenAI’s transition guidance
Some remote browsers remain identifiable. As checked on October 4, 2026, OpenAI documents signed outbound requests for ChatGPT Work’s Cloud browser, verifiable by a CDN, firewall or edge service. A verified signature establishes the requester’s origin; it should not override your site’s authorization rules. Do not trust a signature-related header merely because it is present. OpenAI cloud-browser guidance
Choose controls by the outcome you need
Decline compliant training crawls: maintain operator-specific robots.txt rules. Keep search discovery decisions separate. Cloudflare likewise describes robots.txt compliance as voluntary and says the file does not technically prevent access. Cloudflare documentation
Keep unpublished or paid content away from unauthorized visitors: enforce authentication and authorization before returning the content. Do not list a supposedly secret path in robots.txt and consider it protected; the standard warns that listing paths exposes them publicly. For staging, use the staging access-control workflow. RFC 9309
Authentication does not distinguish a person from an assistant operating that person’s authorized session. It protects the access boundary, not a blanket “no AI” policy.
Restrict identifiable automated traffic: use server-side or edge rules based on verified identities where available. A crawler policy and a request-denial rule are different controls.
Limit extraction or costly actions inside allowed sessions: apply account-aware quotas, rate limits and authorization checks to downloads, exports and write operations. These are proposed safeguards against abuse, not a guaranteed way to distinguish an AI assistant from a person.
Verify the boundary, not just the file
A practical verification sequence:
- Retrieve the public robots.txt response and confirm the intended groups survived CDN processing.
- Test a protected page without credentials. Confirm the content is withheld, not merely hidden in the interface.
- Test an ordinary browser session and an identifiable agent separately. Record the edge action, response status and relevant identity signals; redact credentials and cookies.
- If you rely on request signatures, confirm invalid signatures do not receive trusted-client treatment.
- Keep “identified agent requests” separate from suspected automation. Fast navigation alone does not prove AI involvement.
A successful page request despite Disallow does not necessarily mean the file was parsed incorrectly. It may mean the client did not apply that policy. If access must stop, the decisive check is whether your server refuses to deliver the content—not whether robots.txt asks the client to stay away.