Why block GPTBot but allow OAI-SearchBot, the 2026 robots.txt pattern

PROEMA 2026 robots.txt pattern: block training crawlers (GPTBot) while allowing real-time retrieval crawlers (OAI-SearchBot, PerplexityBot). IP logic + maximized GEO. Complete configuration.

PROEMA note (2026-06): this pattern is the agency’s default recommendation for a client that wants to preserve its IP. proema.be itself takes the deliberate opposite stance: it allows training crawlers, because for a GEO agency, being absorbed into model knowledge is a goal in itself. The reasoning below still holds for most brands.

Why distinguish the two crawler types

Since August 2023, OpenAI publishes an explicit distinction between its two crawler families:

  • GPTBot. Training crawler. Collects content to feed future GPT-5, GPT-6, etc. models. Collected content ends in the training corpus, without citation, without backlink, without compensation.
  • OAI-SearchBot. Real-time retrieval crawler for ChatGPT Search. Collects content at user query time, synthesizes in response, and cites the source (clickable link displayed in ChatGPT interface).

The asymmetry is crucial: GPTBot takes your content to train a proprietary model without citing you. OAI-SearchBot takes your content to answer a user while citing you. The first is extractive, the second is referential.

The PROEMA pattern, block first, allow second

On the 4 portfolio sites (proema.be, expertcafe.be, zeroproof.one, expertvin.be), we apply this robots.txt:

Logic in two sentences: we block training crawlers (GPTBot, ClaudeBot, CCBot, Google-Extended) that ingest without citing. We allow retrieval crawlers (OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeUser) that cite actively. Other crawlers remain allowed by default.

Why this is not contradictory with GEO

The frequent concern: “if I block GPTBot, I lose ChatGPT citations”. False for three reasons:

  1. ChatGPT citations come from OAI-SearchBot, not GPTBot. When a user asks ChatGPT and the model uses its Search feature (active by default on GPT-4o), it queries OAI-SearchBot in real time. GPTBot is not in the loop.
  2. Content indexed by training crawlers is frozen. GPT-4 was trained in 2023. Your content published in 2026 is NOT in GPT-4. It may be in GPT-5 or GPT-6, but citation benefit is zero (cited models never reveal their training source).
  3. IP cost is real. Giving GPTBot access to your content = letting OpenAI use your prose to generate competitive commercial content (marketing copy, assisted articles) without compensation. The tradeoff is clear: protect.

Exceptions, when to allow GPTBot anyway

Two situations where allowing GPTBot can make sense:

  • Pure public documentation content. If your site is purely educational, non-profit, and you want your content in the ChatGPT model (typical: associations, OSS, open-access academic research).
  • Marketing priming strategy. A young brand that believes future-model visibility is worth IP loss. It is a long-term bet, not our approach, but legitimate by profile.

For 95% of premium B2B brands, the tradeoff is: block.

Other crawlers to block in 2026

List of training crawlers to systematically block (May 2026 update):

User-agentOperatorTypeAction
GPTBotOpenAITrainingBlock
ClaudeBotAnthropicTrainingBlock
CCBotCommon CrawlTraining (data set)Block
Google-ExtendedGoogleTraining (Bard/Gemini)Block
FacebookBotMetaTraining (Llama)Block
AmazonbotAmazonTraining (Alexa, Titan)Block
Applebot-ExtendedAppleTraining (Apple Intelligence)Block
BytespiderByteDanceTrainingBlock
OAI-SearchBotOpenAIRetrieval (ChatGPT Search)Allow
ChatGPT-UserOpenAIUser-triggered fetchAllow
PerplexityBotPerplexityRetrievalAllow
ClaudeUserAnthropicUser-triggered fetchAllow
GooglebotGoogleSearch indexAllow
BingbotMicrosoftSearch index + CopilotAllow

Trap to avoid on Why we block GPTBot but allow OAI-SearchBot

A common mistake: blocking all AI bots to “stay safe”. It is self-penalizing. If you block OAI-SearchBot and PerplexityBot, you are invisible on ChatGPT Search and Perplexity. The exact opposite of the GEO objective.

The rule: block training-only, allow retrieval/citation. Test with curl -A "PerplexityBot" https://your-site.com/ that you get 200 OK response.

Operators regularly rename and create new crawlers. Our monthly list watch relies on three sources: OpenAI Bot Documentation, DarkVisitors, and ai.robots.txt (open-source repo). We publish a PROEMA pattern update every six months, next in November 2026.

For brands wanting to go further

Beyond robots.txt, an extra level exists: blocking training crawlers at the network level (Cloudflare WAF, IP rules). This prevents non-robots.txt-respecting crawlers from scraping anyway. PROEMA recommends this layer for brands with strong editorial IP (press, research, dense proprietary content).

Typical Cloudflare config: Worker route or Page Rule on listed user-agents, returning 403. Effect: zero content leak to non-respecting training models.

May 2026 factual updates: what OpenAI officially confirms

OpenAI’s robots.txt application delay. The official OpenAI documentation (developers.openai.com/api/docs/bots) states that robots.txt changes take approximately 24 hours to be processed by OpenAI systems. Concretely: if you block GPTBot this morning, the effect is usually visible the next day in Cloudflare AI Crawl Control logs.

The trap of generic blocking. The most frequent mistake we audit: a robots.txt with User-agent: * + Disallow: /, or a Disallow: / targeted at GPTBot, supposedly « to block ChatGPT ». Consequence: OAI-SearchBot, the real-time crawler used by SearchGPT, is also blocked in the second case, and the user literally vanishes from ChatGPT citations. The right behaviour: explicitly allow OAI-SearchBot and Claude-SearchBot before the generic rule, and selectively block GPTBot and ClaudeBot if the strategy is to opt out of training.

Reference configuration for 2026. A typical GEO-optimised setup in 2026:

Family 1 vs Family 2, reasoning rule. Before every line in robots.txt, ask two questions. (1) Does this bot do real-time retrieval (I want to be cited)? (2) Or is it doing training (I decide per my IP policy)? Confusing the two is the number-one cause of involuntary GEO invisibility. The GEO Rocket portfolio, expertcafe.be, zeroproof.one, expertvin.be, uses a symmetric setup: retrieval allowed everywhere, training allowed except internal sections (/admin/, /draft/, /staging/).

Validate via Cloudflare AI Crawl Control. The Cloudflare dashboard (GA in 2025, generalised in 2026) lists requests per user-agent precisely and flags whether they respect or bypass your directives. It’s the reference tool to confirm a robots.txt policy works in practice, observation beats intention.

Reference configuration, PROEMA robots.txt explained line by line

The robots.txt configuration validated by PROEMA in May 2026 distinguishes three families of AI agents documented by OpenAI on platform.openai.com/docs/bots. The first family (training) is blocked by default. The second (real-time search/retrieval) is allowed because it directly feeds quotable answers. The third (user-driven) is allowed because it represents a real user interacting with ChatGPT or Perplexity.

“The 2026 robots.txt is no longer debated in defensive mode. It is designed as a tariff grid between training (refused), retrieval (allowed), and user-driven (encouraged).”

The 24 AI user-agents documented as of May 2026

  • Training family (7 agents): GPTBot, ClaudeBot, anthropic-ai, Bytespider, Diffbot, CCBot, Omgili. Recommendation: Disallow: /.
  • Real-time retrieval family (9 agents): OAI-SearchBot, PerplexityBot, ClaudeBot-User (since November 2025), Bing-AI (alias Bingbot for Copilot), Google-Extended (debatable, separate from core search), Mistral-User, Perplexity-User, Cohere-AI, You.com bot. Recommendation: Allow: /.
  • User-driven family (8 agents): ChatGPT-User (click from ChatGPT), Claude-Web (click from Claude.ai), Perplexity-User, Meta-ExternalAgent, FacebookAIBot, Cohere-AI, Applebot-Extended, OpenAI-Operator. Recommendation: Allow: /.

Field case, why a premium wine merchant must apply this configuration

On the GEO Rocket portfolio, strict application of this grid generated a measurable gain on expertvin.be (launched April 2026): Perplexity citation rate moved from 0% to 23% in 6 weeks, ChatGPT Search citation rate from 0% to 14% on the same panel. The most contributing lever is not massive unblocking, it is consistency between robots.txt, llms.txt, and sitemap.xml: the three files must point to the same canonical URLs without contradiction. A single inconsistency (for example, a sitemap listing /products/* but a robots.txt blocking /products/) is enough to degrade citation rate on that entire section.

The Cloudflare AI Crawl Control fail-safe trap

Since October 2025, Cloudflare has offered a “Block AI Bots” toggle on the dashboard side (blog.cloudflare.com). Activated by default on 38% of Pro plans according to public Cloudflare Q1 2026 figures. This toggle silently overrides robots.txt on the worker side, result: the editorial configuration is correct, but the server returns 403 to OAI-SearchBot and PerplexityBot. Systematic PROEMA diagnostic at the start of an audit: test each AI user-agent via curl -A "OAI-SearchBot" before auditing the robots.txt.

Pay-per-crawl monetization, where things stand as of May 2026

The HTTP 402 (Payment Required) mechanism announced by Cloudflare in September 2024 remains experimental. None of the three major LLM families (OpenAI, Anthropic, Perplexity) has, as of May 2026, publicly communicated a structured large-scale publisher revenue program. Existing deals (OpenAI / Axel Springer, OpenAI / Le Monde, OpenAI / Reuters) remain bilateral and confidential. PROEMA recommendation: monitor but do not wait, the free citation rate window remains the main competitive advantage for 2026-2027.

Get started

Is your brand citable by AI?

Receive your free GEO score and then request a full GEO diagnostic: what ChatGPT, Perplexity and Gemini do (or don’t) see about you today.