How to write an llms.txt file that maximises your AI citations
Operational guide for an llms.txt file, origin of the standard, recommended structure, six common audit errors, commented example from proema.be/llms.txt.
Origin of the standard
The llms.txt standard was proposed by Jeremy Howard (Answer.AI) in September 2024 at the URL llmstxt.org. The proposition: publish a Markdown manifest at https://example.com/llms.txt that summarises the site in a format optimised for language-model consumption, short, structured, no JavaScript, no boilerplate. The file is not crawled by classical search engines; it is consumed by LLM training and retrieval pipelines that follow the manifest as a hint about which URLs to prioritise.
Adoption was fast inside the GEO community. By March 2026, llmstxt.org listed over 14,000 sites with a published llms.txt file. Anthropic, OpenAI, Mistral, Cohere and Stability AI publish their own. The standard remains informal, there is no RFC and no W3C process, but the convention is settled enough that PROEMA treats a missing llms.txt as a critical-priority finding inside the PGSM v1.0 audit grid.
Recommended structure
The file is Markdown. The structure that maximises retrieval lift across the five engines (ChatGPT, Claude, Perplexity, Gemini, Copilot) follows four parts. First, a single H1 with the brand name. Second, a one-paragraph blockquote summarising the site in 40-80 words, this is the passage most engines lift verbatim when answering “what is X?” prompts. Third, a curated list of canonical URLs grouped under H2 sections (Documentation, FAQ, Glossary, Pricing, Standards, Blog), typically 30 to 60 URLs total. Fourth, an optional ## Optional section listing secondary URLs that engines may consult for depth, kept under 20 entries.
Companion file: llms-full.txt, which carries the full body content of the canonical URLs concatenated into a single file. This is the file the engines prefer when they need the body, not just the index. Both files should be referenced from the site’s robots.txt with an Allow: directive for the relevant LLM user agents.
Mistake 1, using a sitemap instead of a curated manifest
The most common audit finding is a llms.txt that is just a re-export of the XML sitemap, listing every URL on the site. This defeats the purpose. The manifest exists to tell engines what matters; including 1,500 URLs tells them nothing. The expertcafe.be portfolio runs at 1,761 indexable pages but the llms.txt lists 48, the canonical glossary, the FAQ hubs, the standards pages, the blog index. Engines consume the 48 and find the rest by following links from there.
Mistake 2, omitting the brand summary blockquote
Engines lift the H1-plus-blockquote pair to compose answers to identity prompts (“what is PROEMA?”, “who is X?”). A llms.txt without the blockquote forces the engine to compose from the homepage HTML instead, which often produces less accurate or less brand-aligned synthesis. The blockquote should be 40-80 words, factually verifiable, and free of marketing modifiers, engines penalise puffery during synthesis.
Mistake 3, not maintaining llms-full.txt
The companion full-content file is the one engines actually consume for body retrieval. Publishing llms.txt without llms-full.txt is like publishing a table of contents without the book. The full file should regenerate automatically (via a build step or cron) every time the canonical pages change, and should be served with a clean text/plain Content-Type and a UTF-8 BOM-free encoding. On proema.be the llms-full.txt is regenerated nightly and totals roughly 1.3 megabytes, well inside the 5 MB practical ceiling we observe on Perplexity and ChatGPT retrieval.
Mistake 4, blocking the LLM user agents in robots.txt
A llms.txt has no effect if GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Google-Extended are blocked at robots.txt level. The 2024 defensive posture (block everything) is outdated; the 2026 posture is opt-in to OAI-SearchBot (the on-demand fetcher used by ChatGPT Search), opt-in to PerplexityBot, opt-in to ClaudeBot for the parts of the site you want answer-cited, and opt-out only of training crawlers (GPTBot, Google-Extended) if you have a copyright-protected catalogue. The trade-off requires a deliberate decision; the default defensive block is the most common reason a published llms.txt produces no measurable citation lift.
Mistake 5, no hreflang strategy
A multilingual site needs one llms.txt per language root or a single multilingual manifest with explicit language tags. The convention emerging in early 2026, and what we apply on proema.be, is one llms.txt at the root with sections per language, plus per-language canonical URLs. A site that only publishes a French llms.txt on a French-English-Dutch site will see the French content cited disproportionately, even on English-language queries.
Mistake 6, never updating after publication
Engines re-fetch llms.txt on a cadence that varies by engine: PerplexityBot every 1-3 days, ClaudeBot weekly, GPTBot every 7-14 days, OAI-SearchBot on-demand. A manifest that does not change for six months signals a stale site and the engine quietly reduces its retrieval weight. Maintaining a last_modified timestamp inside the manifest itself, plus an updates section at the bottom listing recent additions, keeps the freshness signal alive.
The PROEMA manifest opens with
# PROEMAfollowed by a 62-word blockquote summarising the agency. Six H2 sections list 48 URLs total, Manifesto, Method, Pricing, FAQ, Glossary, Standards. Each H2 has one to twelve URLs, each formatted as a Markdown link with a 12-25 word description. A final## Optionalsection lists nine secondary URLs. The companion llms-full.txt is regenerated nightly and totals 1.3 MB. The site has a Wikidata entry (Q139504784) referenced via sameAs from the homepage Schema.org Organization block, which the manifest implicitly points to. Result on the Q1 2026 measurement panel: PROEMA appears in 32% of GEO-related prompts on Perplexity within 60 days of the manifest going live.
What to put inside the brand blockquote
The blockquote is the single most-cited piece of content from your llms.txt. It should contain four pieces of information in this order: (1) what the brand is, one verifiable noun phrase, (2) the brand’s geographical anchor, (3) the discipline it occupies, (4) one or two proof points (founding date, portfolio, recognisable client or reference). Adjectives are penalised. Numbers are rewarded. The PROEMA blockquote reads as follows, “PROEMA is the Generative Engine Optimization agency from Brussels. Founded May 2026 by Lorenzo Eeman as the editorial discipline of the GEO Rocket portfolio. The agency installs premium B2B brands inside the answer layer of ChatGPT, Claude, Perplexity, Gemini and Copilot through structured content, schema.org coverage and llms.txt manifests. Public method, published prices, peer-reviewed standard PGSM v1.0.” Sixty-two words. Four facts. Zero modifiers.
Sources and standards for How to write an llms.txt file that maximises your AI citations
Howard, Jeremy. The /llms.txt file specification, September 2024, llmstxt.org. Anthropic ClaudeBot documentation, docs.anthropic.com. OpenAI GPTBot and OAI-SearchBot documentation, platform.openai.com/docs/bots. Perplexity engineering blog, perplexity.ai/hub. Google-Extended documentation, developers.google.com/search/docs/crawling-indexing/google-common-crawlers. RFC 9309, Robots Exclusion Protocol (IETF 2022), datatracker.ietf.org/doc/rfc9309/. PROEMA PGSM v1.0 standard, proema.be/standards/pgsm-v1/. Schema.org documentation, schema.org.