Skip to content
LumiRank
The journal
Research10 min read

Every study agrees AI citations are volatile. None of them agree on how volatile.

Published 2026 retention figures for Perplexity range from 44% to 89% depending on who measured. The spread is not a detail — it is the finding, and it changes what an AI visibility report is worth.

By Dmytro Hrysiuk
Six antique brass measuring instruments — a caliper, two balance scales, a ruler, a protractor and a micrometer — standing in a row on a white marble surface, each one measuring an identical small white sphere at a different point on its scale

Executive summary

Ask six published 2026 studies how much of Perplexity's citation set survives from one measurement to the next and you get answers spanning 44% to 89%. For Gemini the published range runs from 11% to 28.1%. These are not different metrics wearing the same name by accident — they are attempts to measure the same underlying property, arriving at numbers that differ by a factor of two, and in some pairings by more.

The industry has responded to this by reporting whichever figure supports the slide. That is the actual problem. Citation volatility is real, well-established and important; the measurement of it is young, unstandardised, and reported with a false precision that nothing in the methodology supports. There is no agreed prompt set, no agreed observation window, no agreed denominator (citations? distinct URLs? distinct domains? brand mentions?) and no agreed treatment of the fact that asking the same question twice returns different answers.

This piece sits inside our work on LLM visibility monitoring, which is where the practice behind it is set out in full.

What does survive across every study, and is worth acting on: a large share of cited URLs are cited once and never again — one study puts it at 73.4%. Cross-engine agreement is very low, with published overlap figures clustering around 8% to 11%. And a meaningful fraction of any brand's citation set rotates every month. Those three findings are robust to the methodological chaos because every study finds them, in the same direction, at similar orders of magnitude. The precise percentages are not, and should stop being quoted as if they were.

What the published figures actually say

Property measuredReported figureEngineSource
Citation retention, 28-day window44%Perplexity5WPR / State of AI Citations 2026
Citation retention, 28-day window11%Gemini5WPR / State of AI Citations 2026
Source retention, day over day28.1%GeminiReported in volatility analyses, 2026
Source retention, day over day~40%ChatGPT, Google AI ModeReported in volatility analyses, 2026
Source retention by source type64%–77%PerplexityVolatility analyses, 2026
Source retention, news/community sources58%–70%Google AI ModeVolatility analyses, 2026
Source retention, editorial/business sources~30%Google AI ModeVolatility analyses, 2026
Monthly citation rotation, mid-sized B2B40%–60%Cross-engineVolatility analyses, 2026
URLs cited exactly once, then never again73.4%Cross-engineVolatility analyses, 2026
Cross-engine domain overlap~11%ChatGPT / Gemini / PerplexityPrinceton GEO research
Cross-engine domain overlap, local queries8%ChatGPT / GeminiSteady Demand, 19 Aug 2026

Read down the Perplexity rows. 44% over 28 days, 64–77% by source type, and "significantly higher" than a ~40% day-over-day baseline in a third framing. Those can all be true simultaneously — a 28-day window should show lower retention than a one-day window, and a per-source-type average is not a per-citation average. That is exactly the point. Nothing in how these figures get quoted signals which window, which denominator, or which unit, and once the number leaves the study it travels alone.

Part 1 — Four ways to measure the same thing and get four answers

The unit. Is a citation a URL, a domain, or a brand mention? A brand cited via five different pages of its own site is one domain, five URLs, and — depending on the parser — one mention or five. Retention measured on domains will always look far more stable than retention measured on URLs, because a site can lose every individual page from the citation set while keeping its domain presence intact. The 73.4% single-appearance figure is a URL-level statistic. Quoting it next to a domain-level retention rate compares two different things.

The window. Day-over-day, week-over-week and 28-day retention are different quantities and decay is not linear. A study reporting 40% day-over-day and a study reporting 44% over 28 days are not in conflict; they are barely in conversation.

The prompt set. Every study picks its own questions. Prompt sets skewed toward commercial queries, toward one vertical, or toward one country will produce systematically different source mixes. Some studies filter to prompts that return stable results across repeated runs before measuring stability, which is close to circular and is rarely flagged in the summary.

The run-to-run problem. Ask any of these systems the same question twice and the answers differ. That variance is inside every measurement, and almost no published figure reports a confidence interval. A "44% retention" that carries an unstated ±15 is a different claim from one that carries ±2, and the reader is given no way to tell which they have.

Part 2 — Who is publishing, and what that changes

Most of this research is vendor research. The companies measuring AI citation volatility mostly sell AI visibility monitoring, and a finding that citations are volatile, unpredictable and in need of continuous tracking is a finding that sells the product. That does not make the numbers wrong. It does mean the incentive runs one way, and the incentive plus the absence of a standard methodology is a combination that should slow anyone down.

As of 31 August 2026 there is one count nobody sells. Search Console's generative AI report gives every property a first-party impression total for AI Overviews and AI Mode, which settles none of the arguments above — it carries no clicks and no queries — but it is the first figure in this field measured from inside the system rather than sampled from outside it. It reached general availability under a regulatory order, without ever entering Google's ranking-update log.

The exception in the table above is the Princeton GEO work, which is academic, peer-reviewed, and reports a cross-engine overlap figure of roughly 11%. It lands close to Steady Demand's 8% for local queries despite completely different methods and scopes. Two independent estimates of low cross-engine agreement, arriving at single-digit-to-low-double-digit overlap from different directions, is the strongest result in this whole literature — and it is the one least often quoted, because it does not produce a number anyone can put on a dashboard.

This is a pattern worth naming, because it keeps recurring in this field. The most-cited AI citation study of the previous cycle turned out to have run on GPT-3.5 against five fabricated search results, which is documented at length in The 40% Citation Boost in the Famous GEO Study Was Simulated. The industry's appetite for a quotable number consistently outruns its appetite for checking where the number came from.

Part 3 — What survives the chaos

Three findings are robust enough to plan around, because every study finds them in the same direction.

Citation sets churn heavily. Whether the monthly rotation is 40% or 60%, a brand's AI citation footprint is not a position it holds. It is a distribution it occupies, resampled continuously. Anyone reporting AI visibility as a rank is describing a property the system does not have.

A single URL rarely repeats. The 73.4% figure may be soft, but URL-level churn far exceeding domain-level churn is consistent across sources. The practical reading: engines are re-picking a passage for each query, not remembering that your page was good last week. That fits everything else known about how retrieval works, including Google's own documented fan-out behaviour and the gap between what it discloses about retrieval and what it discloses about selection, covered in Google Explained How AI Overviews Retrieve Content.

The engines barely agree with each other. Whether overlap is 8% or 11%, it is small. There is no single "AI visibility" to win. There are five or six largely independent visibility problems that happen to be measured with the same vocabulary, and a strategy tuned to one of them transfers poorly to the rest.

Part 4 — What a defensible AI visibility report looks like

The fix is not more precision. It is honest reporting of the precision that exists.

State the window and the unit on every figure, in the figure. "Cited in 18% of runs" is meaningless; "cited in 18% of 200 runs across 40 prompts, ChatGPT Search, week of 8 September" is a claim someone can check and reproduce. Run every prompt multiple times and report the range, not the mean — a metric whose whole subject is volatility should not be summarised by a single number. Track each engine separately and never average across them, because at 8–11% overlap the average describes nothing real. Compare against your own prior measurement taken the same way, not against a published industry benchmark that used a different prompt set, a different window and a different denominator.

And measure often enough to see a step change. A domain-wide crawler block took one major platform's ChatGPT citation share down 86% in six days, which monthly reporting would have shown as a single ugly datapoint a month later with no way to identify the cause. That episode is worked through in Reddit Blocked the Crawlers.

Counter-argument, taken seriously

Demanding methodological rigour from a two-year-old field may be the wrong ask. Early web analytics were a mess too — competing definitions of a "visit," server logs against page tags, numbers that differed by 30% between vendors measuring the same site. The field converged, slowly, because practitioners used the tools anyway and the disagreements became intolerable. Refusing to measure until the methodology settles would mean flying blind through the exact period when the ground is moving fastest.

That objection is largely right, and it is why this piece argues for stated error bars rather than abstention. The narrower and harder version of the critique stands, though: web analytics vendors were not simultaneously the primary publishers of the research establishing that web analytics were necessary. Here they mostly are. The remedy is not to stop measuring but to weight academic and independently-replicated findings above vendor benchmarks, and to treat any number without a disclosed prompt set and sample size as directional at best.

What everyone is missing

The spread between studies is more useful than any individual study. A factor-of-two disagreement on Perplexity retention tells you the real uncertainty in the field. That is a genuinely valuable number, and it is available only by reading across the literature rather than citing one paper from it.

Almost nobody reports sample size, and almost nobody is asked to. Across the retention figures gathered here, disclosed sample sizes are the exception. In any other measurement discipline that would disqualify a figure from being quoted. In this one it does not slow it down at all.

"AI visibility" is a category error as a single metric. At 8–11% cross-engine overlap, ChatGPT visibility and Gemini visibility are about as related as Google rankings and Yelp rankings. Reporting them under one heading invites exactly the wrong strategic response: one programme, averaged, tuned to nothing.

Volatility may be a feature of the query, not the brand. None of the published work separates how much of the churn comes from the engine resampling versus from the prompt set containing genuinely unstable questions. A prompt with one obvious answer will be stable everywhere; a prompt with twenty defensible answers will look volatile no matter who is cited. Until someone stratifies by query determinism, a portion of every published volatility figure is measuring the questions rather than the answers.

Future predictions

  • A standard emerges from an academic or consortium source, not a vendor. The incentive to standardise sits with the people who do not sell dashboards.
  • Confidence intervals become table stakes within a year. The first vendor to publish ranges instead of point estimates will use it as a credibility differentiator, and the rest will follow.
  • Cross-engine overlap stays low. Nothing in the current architecture pushes these systems toward agreement; they have different indexes, different grounding sources and different retrieval layers.
  • Someone publishes query-determinism stratification and a chunk of the reported volatility disappears into it. Expect the "real" engine-driven churn to be lower than current headline figures once prompt instability is separated out.

Practical takeaways

  1. Never quote an AI citation figure without its window, unit and sample size. If those are not in the source, the figure is directional and should be presented that way.
  2. Run every prompt at least five times and report the range. A single run of a stochastic system is an anecdote.
  3. Report per engine. Never average across engines. At 8–11% overlap, the average is a number describing nothing.
  4. Benchmark against your own prior measurement, taken identically. Cross-vendor benchmarks compare incompatible methods.
  5. Weight academic and independently-replicated findings above vendor benchmarks when the two disagree, which is often.
  6. Measure weekly at minimum. Step changes in this space happen inside a week. See LLM Visibility Monitoring.
  7. Optimise for being a durable, resolvable entity rather than for holding a citation. URL-level churn is high and domain-level presence is stickier, which points the work at entity architecture rather than at individual pages. See Entity & Knowledge Architecture.

Read the rest of the journal.

Key takeaways
  • Published 2026 retention figures for the same engine differ by roughly a factor of two — Perplexity appears at 44%, at 64%–77%, and as "significantly higher" than a ~40% baseline across three separate framings.
  • The disagreement traces to four unstandardised choices: the unit measured (URL, domain, brand mention), the observation window, the prompt set, and whether prompts are pre-filtered for stability.
  • Three findings are robust across every study: citation sets churn heavily (40%–60% monthly rotation), individual URLs rarely repeat (one analysis: 73.4% appear exactly once), and cross-engine agreement is very low (8%–11% domain overlap).
  • Most of this research is published by vendors selling AI visibility monitoring. The academic figure (Princeton, ~11% overlap) and an independent local-query study (8%) agree closely despite different methods, which makes low cross-engine overlap the best-supported result in the literature.
  • "AI visibility" as a single averaged metric is a category error at 8%–11% overlap. Report per engine.
  • No published volatility figure separates engine-driven churn from prompt instability, so an unknown share of every headline number is measuring the questions rather than the answers.
  • A defensible report states window, unit and sample size on every figure, runs each prompt multiple times, reports ranges rather than means, and benchmarks against its own prior measurement.

Frequently asked

How stable is an AI citation once you have one?
Less stable than a search ranking, and nobody can tell you precisely how much less. Published 2026 retention figures for the same engine differ by roughly a factor of two — Perplexity appears at 44% over a 28-day window in one study and at 64%–77% by source type in another. What is consistent across studies is that URL-level citations churn heavily, with one analysis finding 73.4% of cited URLs appear exactly once and never again.
Why do different AI visibility studies report such different numbers?
Because there is no shared methodology. Studies differ on the unit measured (URL, domain, or brand mention), the observation window (day-over-day against 28-day), the prompt set, and whether they filter for prompts that already return stable results. All four choices move the headline number substantially, and none of them travel with the figure once it is quoted.
Do the major AI engines cite the same sources?
Barely. Princeton's GEO research puts cross-engine domain overlap at roughly 11% across ChatGPT, Gemini and Perplexity. A separate August 2026 study of 14,472 citations from 1,487 local queries found ChatGPT and Gemini citing the same domains only 8% of the time. Two independent methods, both landing in single digits to low double digits.
Is AI visibility research trustworthy?
Treat it as directional. Most of it is published by companies that sell AI visibility monitoring, which is a real conflict even when the work is honest. Weight academic and independently-replicated results higher, insist on disclosed sample sizes and prompt sets, and be suspicious of any figure quoted to two decimal places without an error bar.
How often should we measure AI citations?
Weekly at minimum. A crawler-access change took one major platform's ChatGPT citation share down 86% in six days in August 2026. Monthly reporting would have recorded that as one bad datapoint with no recoverable cause.
From the journal

Findings like this are only useful against your own numbers. That starts with a baseline: