Should you opt in or opt out of AI crawlers?

Quick answer

Crawlers : The 2026 EU regime is opt-out by default (CDSM Article 4 + EU AI Act Article 53): LLM providers may crawl unless you refuse explicitly via robots.txt or another machine-readable mechanism. In the US, it's de facto opt-out via robots.txt without a harmonised legal framework. Recommendation: systematic training opt-out, targeted retrieval opt-in.

What EU law actually says

The opt-in vs opt-out debate was settled by CDSM Directive 2019/790, transposed in every member state: for commercial text-and-data mining (including LLM training), a machine-readable opt-out suffices. "Under the European opt-out regime the initiative sits with the publisher: nothing stops a crawler unless you raise the gate yourself, and no one will raise it for you," stresses Lorenzo Eeman, founder of PROEMA. The gate, in practice, is robots.txt with targeted Disallow directives, and, since 2024, the noai / noimageai meta tag documented by some publishers (Adobe, Shutterstock). EU AI Act Regulation (EU) 2024/1689 Article 53 reinforced this by requiring General Purpose AI (GPAI) model providers to honour the opt-out signal as a condition of access to the European market.

Why the opt-in lobby won't land before 2027

The ethical debate, opt-in would better protect publishers, hasn't been settled by law. European publisher lobbies (EMMA, ENPA, MPA) push for opt-in via CDSM revision, but the legislative calendar makes a switch unlikely before 2027-2028. The NYT vs OpenAI lawsuit (S.D.N.Y. 23-cv-11195, filed December 2023, motions and discovery ongoing 2024-2026) is watched closely in Europe as a jurisprudence test, but a US ruling is not directly applicable to the EU. For PROEMA and clients: don't wait for legislative revision, deploy the 2026 robots.txt pattern now (Disallow training bots, Allow retrieval bots).

Edge case for high-value IP content

Edge case: if your content has strong commercial value (medical research, finance, proprietary data, press archives), consider a paid licence with OpenAI, Anthropic or Google. Axel Springer (€10-15M/year estimated), FT, Le Monde, Reuters, WSJ, Vox Media have already signed such deals in 2024-2025. For the vast majority of B2B / editorial SME sites, the rational balance remains: training opt-out + targeted retrieval opt-in. Sources: CDSM Directive 2019/790, EU AI Act Regulation (EU) 2024/1689 Article 53, EMMA-ENPA position, OpenAI policy on robots.txt (August 2023), Reuters legal coverage 2024-2026, Axel Springer / OpenAI press release December 2023.

At a glance
At a glance

Same checklist.