Why do we talk about tokens in LLMs?
Tokens : A token is the elementary unit an LLM manipulates. It represents roughly 0.75 word in English and 0.7 word in French, a word fragment, a short word, or even a punctuation mark. All API billing, context limits and model performance are measured in tokens, never words or characters.
Why do we talk about tokens in LLMs, in numbers?
Why tokens, not words? "LLMs work in tokens rather than words for two reasons, starting with (1) covering all languages: no vocabulary of whole words can fit more than a hundred languages," explains Lorenzo Eeman, founder of PROEMA. The tokenizer fragments into frequent sub-units. (2) Handle unknown words, a rare brand name becomes a sequence of known tokens. Practical consequences: (a) Cost, English cheapest (1 word = 1 token); French ~30% more; Japanese or German ~50% more. (b) Limits, 128k token window ≈ 95k EN words or ~85k FR words. (c) Performance, on languages with many tokens per word, the model "reasons" with effectively stretched context. For GEO, your key contents (key pages, "About", "Method") should be readable in under 2,000 tokens to fit the LLM's short reasoning window.
Technical detail moving the LLM needle on Why do we talk about tokens in LLMs
Three often-forgotten fragments tip citation outcomes. (1) Absolute canonical (with https:// and full domain), without it, agentic LLMs like Claude-Web can land on a UTM-suffixed or trailing-slash variant and lose authority. (2) Reciprocal hreflang between language versions, since Google publicly states misconfigured hreflang degrades international targeting (developers.google.com/search). (3) JSON-LD Schema.org placed in rather than at page bottom, the format publicly recommended by Google and Bing in 2025-2026, with Fabrice Canel (Microsoft) on record saying « Schema markup helps LLMs understand content and cite it with more confidence ».
How to audit Why do we talk about tokens in LLMs in under an hour
Three tools cover any page. (1) Google Rich Results Test to validate Schema.org and surface JSON-LD errors. (2) Schema.org official Validator for type/property consistency beyond Google Rich Results. (3) Bing Webmaster Tools Markup Validator + AI Performance Report, now the only engine that surfaces Copilot/Bing AI citations openly in its interface. Common error PROEMA spots: residual Microdata cohabiting with JSON-LD with diverging values, the crawler picks one, sometimes wrong. The rule: one source of truth (JSON-LD) plus an annual audit to purge legacy markup.
30-minute self-audit on Why do we talk about tokens in LLMs
Open your home page in a fresh tab, hit F12 (DevTools) → Elements tab → search « application/ld+json ». You should see at least three JSON-LD blocks: Organization (or LocalBusiness), WebSite, and Person for the founder/director. Missing one? That's a citation-rate gap. Same drill on a content page: FAQPage + Article + Author Person with sameAs. Two minutes per page, thirty minutes for the top ten pages of the site. This single audit surfaces 80 % of the Schema.org issues PROEMA finds in initial diagnostics.
| Language | Words per 1,000 tokens | Relative cost |
|---|---|---|
| English | ~750 | 1.0× |
| Spanish | ~700 | 1.07× |
| French | ~700 | 1.07× |
| German | ~650 | 1.15× |
| Dutch | ~660 | 1.14× |
| Italian | ~700 | 1.07× |
| Chinese | ~500 | 1.5× |
| Japanese | ~450 | 1.67× |