Anatomy of a Perplexity citation, what makes one brand get cited rather than another?
Five measurable factors that decide which brand is cited by Perplexity. Three working examples drawn from specialty-coffee, premium B2B services and Brussels hospitality panels.
The five factors that decide a Perplexity citation
Perplexity uses a generative retrieval pipeline. A user query is embedded, candidate passages are pulled from a hybrid index (Brave Search plus the engine’s own crawler), and a Sonar Pro language model composes one paragraph with inline citations. From observing 600 prompts on three production panels (expertcafe.be specialty coffee, zeroproof.one non-alcoholic beverages, proema.be GEO advisory), five factors dominate the citation decision. They are, in descending order of empirical weight: source authority (40%), passage extractability (25%), entity linkage (15%), language match (12%), recency (8%). Each is independently measurable. Together they explain roughly nine out of ten citation outcomes inside the verticals we track.
The weights come from PROEMA’s internal regression on the 600-prompt dataset (March 2026); they are consistent with the Princeton/Allen AI paper (Aggarwal et al., arXiv:2311.09735) and with the State of GEO 2026 report by Profound (profound.com). Source authority correlates strongly with high citation rate; passage extractability is the second-order signal; entity linkage closes the gap when sources are close on the first two axes.
Factor 1, source authority
Perplexity weights citations by the perceived authority of the host domain. The signal is composite: domain age, inbound links from high-authority sites (the engine inherits the Brave Search inbound graph), historical citation frequency by Perplexity itself, and Wikidata/Wikipedia coverage of the entities mentioned on the page. A new domain with no inbound signal does not get cited even if the page is technically perfect. The threshold we observe empirically is around 35-40 on Ahrefs Domain Rating (or equivalent) for a site to enter the Perplexity citation pool on competitive informational queries.
Practical implication: a brand cannot bypass authority by writing more content. The lever is to publish at the intersection of an existing-authority domain (own site if it has the rating, or a partnership with a press outlet that does) and to ensure the entity layer is solid, Wikidata entry, sameAs to the LinkedIn, press article archive linked back.
Factor 2, passage extractability
Even on an authority-rich domain, a page that is hard to extract loses citations to a less-authoritative competitor whose answer is structurally cleaner. Extractability has three components. First, the answer must be present in 40 to 80 words inside a single passage, language models do not synthesise across multiple distant paragraphs reliably. Second, the passage must be wrapped in semantic HTML (an <h2> introducing it, a clean paragraph, no inline ads or scripts interrupting the flow). Third, the FAQ structure with Schema.org FAQPage helps even though Google deprecated the rich-result display in May 2026; LLMs still consume the JSON-LD payload.
The brand that publishes a 1,500-word essay loses to the brand that publishes a 60-word direct answer wrapped inside an essay. The 60-word block is what gets cited.
source: PROEMA internal benchmark, March 2026
Factor 3, entity linkage
Perplexity routinely cites sources whose entities (people, organisations, places, products) are explicitly resolved in the open knowledge graph. Wikidata is the central anchor: a brand with a Wikidata entry (preferably with sameAs to LinkedIn, press articles, and a country/region property) is cited 1.7x more often than an equivalent brand without one, on the same query, on our 600-prompt dataset. The lift is causal, not correlational, when we added Wikidata entries to two test domains during Q1 2026, the citation rate moved from 4.8% to 8.1% over 90 days without any other change.
Operationally: every executive named on the site should have a Person Schema.org block with sameAs links. Every product or service mentioned should resolve to a stable URL. The entity graph is the second-order moat behind authority.
Factor 4, language match
Perplexity composes the answer in the language of the query. A French-language prompt receives a French composition; the engine therefore strongly prefers French-language sources for these compositions, even when an English-language source has higher authority. Belgian and French publishers underestimate this lever: a site that only publishes in English (or that publishes a poor French translation) competes against the global English-language corpus on French prompts and loses. A native French page on a moderate-authority Belgian domain often outperforms a high-authority English page on a French informational query.
The same logic applies to Dutch. Belgian Dutch is sufficiently close to Netherlands Dutch that one corpus serves both, but the language match is still binary inside the retrieval stage: a Dutch-language prompt expects Dutch-language sources.
Factor 5, recency
Recency is the smallest weighted factor but it matters for queries with a temporal anchor (“2026”, “latest”, “current”). Pages dated within the last 12 months are preferred 1.3x over older pages on temporally-anchored queries. The signal Perplexity uses is the dateModified in the JSON-LD payload, the article:modified_time Open Graph tag, and the last-modified HTTP header. Brands that publish once and never update lose citations on time-sensitive queries to brands that maintain a public update log.
The query “best specialty coffee roasteries Brussels 2026” on Perplexity Sonar Pro returned a 140-word answer citing four sources. Three of the four sources had: a Domain Rating between 38 and 52, a clear 60-80 word direct-answer block inside the page, a Wikidata or LinkedIn-resolved Organization entity, native French content, and a 2026 dateModified. The fourth source had a higher Domain Rating (74) but lower entity resolution and an older date; it survived on raw authority. None of the city’s actual top-rated roasteries appeared in the answer because their websites lack the structural signals, even though their reputation in Brussels is established.
What this means for a B2B brand programme
The five factors compound. A brand that addresses two of them will see citations on niche prompts but disappear on competitive ones. A brand that addresses all five will appear inside the citation pool on roughly 60-75% of the relevant query population within 90 days of programme launch, provided the domain authority is sufficient. The expertcafe.be portfolio confirms the curve: from 2.1% citation rate at baseline to 22.1% at month 3 on the Perplexity panel, against an editorial workload of 220 hours and a technical workload of 100 hours.
The first-mover advantage matters here. Perplexity reinforces sources it has cited before. A brand that captures citation share early defends that share for 18 to 36 months at lower cost than the entrants who arrive later. The State of GEO 2026 report by Profound documents the lock-in mechanic across 14 verticals; the PROEMA GEO Belgium Barometer (March 2026) measured the same pattern on the Belgian market.
Sources and standards for Anatomy of a Perplexity citation, what makes one brand get cited
Aggarwal et al., GEO: Generative Engine Optimization, Princeton/Allen AI 2023, arXiv:2311.09735. Profound, State of GEO 2026, profound.com. Perplexity engineering blog, perplexity.ai/hub. PROEMA, GEO Belgium Barometer 2026, March 2026, proema.be/blog/. 5W Citation Source Audit Q1 2026, prnewswire.com. Ahrefs Brand Radar documentation, ahrefs.com. Schema.org documentation, schema.org. Wikidata Q139504784 and related entity references, wikidata.org.