The 2026 robots.txt pattern: Allow search, Disallow training, step by step
The robots.txt pattern that took hold in 2026 splits two AI bot uses: live search (allow) and training (disallow). A how-to.
PROEMA note (2026-06): this pattern is the agency’s default recommendation for a client that wants to preserve its IP. proema.be itself takes the deliberate opposite stance: it allows training crawlers, because for a GEO agency, being absorbed into model knowledge is a goal in itself. The reasoning below still holds for most brands.
1. The stake: two uses, two regimes
The same domain can be read by AI bots in two radically different ways. The first use is training, where content is ingested to feed an LLM’s foundation model without citation, link, or revenue to the source brand. The second is live search, where the bot fetches content in real time to answer a user query, with explicit citation and back-link.
The 2026 pattern allows search, which brings visibility, and blocks training, which brings none. That is the arbitrage carried by GEO Rocket portfolios since April 2026 and recommended by the PGSM v1.0 standard.
Allow search, disallow training. This two-step grammar became the implicit B2B standard by June 2026.
2. The bots to name explicitly
RFC 9309 governs robots.txt grammar since 2022. Main bots to allow for search: OAI-SearchBot (OpenAI live search), PerplexityBot (Perplexity), ClaudeBot (Anthropic in citation mode), Google-Extended (Google AI Overviews and Gemini live), Bingbot (Bing Copilot). Main bots to block for training: GPTBot (OpenAI training), CCBot (Common Crawl), anthropic-ai (Anthropic training).
Three bots are ambiguous and warrant a brand-by-brand decision: Applebot-Extended (Apple Intelligence training), facebookexternalhit (Meta training), bytespider (ByteDance). PROEMA standard recommends blocking by default unless explicit commercial intent toward those ecosystems.
3. The 2026 robots.txt template
Here is the template carried by the three GEO Rocket portfolio sites. It is public, verifiable at expertcafe.be/robots.txt, zeroproof.one/robots.txt, expertvin.be/robots.txt.
| Bot | Policy | Rationale |
|---|---|---|
| OAI-SearchBot | Allow | ChatGPT live search, citation source |
| PerplexityBot | Allow | Direct Perplexity citations |
| Google-Extended | Allow | AI Overviews and Gemini search |
| Bingbot | Allow | Copilot and Bing AI Performance |
| ClaudeBot | Allow | Claude 3.5+ citations |
| GPTBot | Disallow | OpenAI training, zero revenue |
| CCBot | Disallow | Third-party training, zero control |
| anthropic-ai | Disallow | Anthropic training distinct from ClaudeBot |
4. The wildcard trap
The most common mistake, observed on nearly 40% of French-speaking B2B sites audited in Q1 2026, is using User-agent: * Disallow: /. That single line blocks everything, search and training alike. Result: the site disappears both from LLMs and from Google.
Standard practice writes each allowed bot explicitly and avoids wildcards. RFC 9309 specifies that bots evaluate rules in order, stopping at the first match. A misordered policy silently blocks the right bots.
5. How to verify the policy works
Three tools verify in under ten minutes. Google Search Console’s robots.txt tester simulates Googlebot evaluation. The Cloudflare AI Crawl Control report logs every AI bot visit and its treatment in real time. The curl command with the targeted bot’s User-Agent header validates the expected HTTP response.
Golden rule: test every robots.txt deployment on these three tools before publication, and keep Cloudflare logs for three months to verify no bot changes behaviour without the brand noticing.
robots.txt in 2026 is no longer a hidden technical file. It is the opening line of a brand’s GEO policy, read by bots and auditable by competitors.
source: PGSM v1.0 standard · robots.txt section
The PROEMA GEO diagnostic includes a full 15-point audit of the robots.txt file. Read in five minutes, recommendations directly copyable into the file.