Executive summary
Ask six published 2026 studies how much of Perplexity's citation set survives from one measurement to the next and you get answers spanning 44% to 89%. For Gemini the published range runs from 11% to 28.1%. These are not different metrics wearing the same name by accident — they are attempts to measure the same underlying property, arriving at numbers that differ by a factor of two, and in some pairings by more.
The industry has responded to this by reporting whichever figure supports the slide. That is the actual problem. Citation volatility is real, well-established and important; the measurement of it is young, unstandardised, and reported with a false precision that nothing in the methodology supports. There is no agreed prompt set, no agreed observation window, no agreed denominator (citations? distinct URLs? distinct domains? brand mentions?) and no agreed treatment of the fact that asking the same question twice returns different answers.
What does survive across every study, and is worth acting on: a large share of cited URLs are cited once and never again — one study puts it at 73.4%. Cross-engine agreement is very low, with published overlap figures clustering around 8% to 11%. And a meaningful fraction of any brand's citation set rotates every month. Those three findings are robust to the methodological chaos because every study finds them, in the same direction, at similar orders of magnitude. The precise percentages are not, and should stop being quoted as if they were.
What the published figures actually say
| Property measured | Reported figure | Engine | Source |
|---|
| Citation retention, 28-day window | 44% | Perplexity | 5WPR / State of AI Citations 2026 |
| Citation retention, 28-day window | 11% | Gemini | 5WPR / State of AI Citations 2026 |
| Source retention, day over day | 28.1% | Gemini | Reported in volatility analyses, 2026 |
| Source retention, day over day | ~40% | ChatGPT, Google AI Mode | Reported in volatility analyses, 2026 |
| Source retention by source type | 64%–77% | Perplexity | Volatility analyses, 2026 |
| Source retention, news/community sources | 58%–70% | Google AI Mode | Volatility analyses, 2026 |
| Source retention, editorial/business sources | ~30% | Google AI Mode | Volatility analyses, 2026 |
| Monthly citation rotation, mid-sized B2B | 40%–60% | Cross-engine | Volatility analyses, 2026 |
| URLs cited exactly once, then never again | 73.4% | Cross-engine | Volatility analyses, 2026 |
| Cross-engine domain overlap | ~11% | ChatGPT / Gemini / Perplexity | Princeton GEO research |
| Cross-engine domain overlap, local queries | 8% | ChatGPT / Gemini | Steady Demand, 19 Aug 2026 |
Read down the Perplexity rows. 44% over 28 days, 64–77% by source type, and "significantly higher" than a ~40% day-over-day baseline in a third framing. Those can all be true simultaneously — a 28-day window should show lower retention than a one-day window, and a per-source-type average is not a per-citation average. That is exactly the point. Nothing in how these figures get quoted signals which window, which denominator, or which unit, and once the number leaves the study it travels alone.
Part 1 — Four ways to measure the same thing and get four answers
The unit. Is a citation a URL, a domain, or a brand mention? A brand cited via five different pages of its own site is one domain, five URLs, and — depending on the parser — one mention or five. Retention measured on domains will always look far more stable than retention measured on URLs, because a site can lose every individual page from the citation set while keeping its domain presence intact. The 73.4% single-appearance figure is a URL-level statistic. Quoting it next to a domain-level retention rate compares two different things.
The window. Day-over-day, week-over-week and 28-day retention are different quantities and decay is not linear. A study reporting 40% day-over-day and a study reporting 44% over 28 days are not in conflict; they are barely in conversation.
The prompt set. Every study picks its own questions. Prompt sets skewed toward commercial queries, toward one vertical, or toward one country will produce systematically different source mixes. Some studies filter to prompts that return stable results across repeated runs before measuring stability, which is close to circular and is rarely flagged in the summary.
The run-to-run problem. Ask any of these systems the same question twice and the answers differ. That variance is inside every measurement, and almost no published figure reports a confidence interval. A "44% retention" that carries an unstated ±15 is a different claim from one that carries ±2, and the reader is given no way to tell which they have.
Part 2 — Who is publishing, and what that changes
Most of this research is vendor research. The companies measuring AI citation volatility mostly sell AI visibility monitoring, and a finding that citations are volatile, unpredictable and in need of continuous tracking is a finding that sells the product. That does not make the numbers wrong. It does mean the incentive runs one way, and the incentive plus the absence of a standard methodology is a combination that should slow anyone down.
As of 31 August 2026 there is one count nobody sells. Search Console's generative AI report gives every property a first-party impression total for AI Overviews and AI Mode, which settles none of the arguments above — it carries no clicks and no queries — but it is the first figure in this field measured from inside the system rather than sampled from outside it. It reached general availability under a regulatory order, without ever entering Google's ranking-update log.
The exception in the table above is the Princeton GEO work, which is academic, peer-reviewed, and reports a cross-engine overlap figure of roughly 11%. It lands close to Steady Demand's 8% for local queries despite completely different methods and scopes. Two independent estimates of low cross-engine agreement, arriving at single-digit-to-low-double-digit overlap from different directions, is the strongest result in this whole literature — and it is the one least often quoted, because it does not produce a number anyone can put on a dashboard.
This is a pattern worth naming, because it keeps recurring in this field. The most-cited AI citation study of the previous cycle turned out to have run on GPT-3.5 against five fabricated search results, which is documented at length in The 40% Citation Boost in the Famous GEO Study Was Simulated. The industry's appetite for a quotable number consistently outruns its appetite for checking where the number came from.
Part 3 — What survives the chaos
Three findings are robust enough to plan around, because every study finds them in the same direction.
Citation sets churn heavily. Whether the monthly rotation is 40% or 60%, a brand's AI citation footprint is not a position it holds. It is a distribution it occupies, resampled continuously. Anyone reporting AI visibility as a rank is describing a property the system does not have.
A single URL rarely repeats. The 73.4% figure may be soft, but URL-level churn far exceeding domain-level churn is consistent across sources. The practical reading: engines are re-picking a passage for each query, not remembering that your page was good last week. That fits everything else known about how retrieval works, including Google's own documented fan-out behaviour and the gap between what it discloses about retrieval and what it discloses about selection, covered in Google Explained How AI Overviews Retrieve Content.
The engines barely agree with each other. Whether overlap is 8% or 11%, it is small. There is no single "AI visibility" to win. There are five or six largely independent visibility problems that happen to be measured with the same vocabulary, and a strategy tuned to one of them transfers poorly to the rest.
Part 4 — What a defensible AI visibility report looks like
The fix is not more precision. It is honest reporting of the precision that exists.
State the window and the unit on every figure, in the figure. "Cited in 18% of runs" is meaningless; "cited in 18% of 200 runs across 40 prompts, ChatGPT Search, week of 8 September" is a claim someone can check and reproduce. Run every prompt multiple times and report the range, not the mean — a metric whose whole subject is volatility should not be summarised by a single number. Track each engine separately and never average across them, because at 8–11% overlap the average describes nothing real. Compare against your own prior measurement taken the same way, not against a published industry benchmark that used a different prompt set, a different window and a different denominator.
And measure often enough to see a step change. A domain-wide crawler block took one major platform's ChatGPT citation share down 86% in six days, which monthly reporting would have shown as a single ugly datapoint a month later with no way to identify the cause. That episode is worked through in Reddit Blocked the Crawlers.
Counter-argument, taken seriously
Demanding methodological rigour from a two-year-old field may be the wrong ask. Early web analytics were a mess too — competing definitions of a "visit," server logs against page tags, numbers that differed by 30% between vendors measuring the same site. The field converged, slowly, because practitioners used the tools anyway and the disagreements became intolerable. Refusing to measure until the methodology settles would mean flying blind through the exact period when the ground is moving fastest.
That objection is largely right, and it is why this piece argues for stated error bars rather than abstention. The narrower and harder version of the critique stands, though: web analytics vendors were not simultaneously the primary publishers of the research establishing that web analytics were necessary. Here they mostly are. The remedy is not to stop measuring but to weight academic and independently-replicated findings above vendor benchmarks, and to treat any number without a disclosed prompt set and sample size as directional at best.
What everyone is missing
The spread between studies is more useful than any individual study. A factor-of-two disagreement on Perplexity retention tells you the real uncertainty in the field. That is a genuinely valuable number, and it is available only by reading across the literature rather than citing one paper from it.
Almost nobody reports sample size, and almost nobody is asked to. Across the retention figures gathered here, disclosed sample sizes are the exception. In any other measurement discipline that would disqualify a figure from being quoted. In this one it does not slow it down at all.
"AI visibility" is a category error as a single metric. At 8–11% cross-engine overlap, ChatGPT visibility and Gemini visibility are about as related as Google rankings and Yelp rankings. Reporting them under one heading invites exactly the wrong strategic response: one programme, averaged, tuned to nothing.
Volatility may be a feature of the query, not the brand. None of the published work separates how much of the churn comes from the engine resampling versus from the prompt set containing genuinely unstable questions. A prompt with one obvious answer will be stable everywhere; a prompt with twenty defensible answers will look volatile no matter who is cited. Until someone stratifies by query determinism, a portion of every published volatility figure is measuring the questions rather than the answers.
Future predictions
- A standard emerges from an academic or consortium source, not a vendor. The incentive to standardise sits with the people who do not sell dashboards.
- Confidence intervals become table stakes within a year. The first vendor to publish ranges instead of point estimates will use it as a credibility differentiator, and the rest will follow.
- Cross-engine overlap stays low. Nothing in the current architecture pushes these systems toward agreement; they have different indexes, different grounding sources and different retrieval layers.
- Someone publishes query-determinism stratification and a chunk of the reported volatility disappears into it. Expect the "real" engine-driven churn to be lower than current headline figures once prompt instability is separated out.
Practical takeaways
- Never quote an AI citation figure without its window, unit and sample size. If those are not in the source, the figure is directional and should be presented that way.
- Run every prompt at least five times and report the range. A single run of a stochastic system is an anecdote.
- Report per engine. Never average across engines. At 8–11% overlap, the average is a number describing nothing.
- Benchmark against your own prior measurement, taken identically. Cross-vendor benchmarks compare incompatible methods.
- Weight academic and independently-replicated findings above vendor benchmarks when the two disagree, which is often.
- Measure weekly at minimum. Step changes in this space happen inside a week. See LLM Visibility Monitoring.
- Optimise for being a durable, resolvable entity rather than for holding a citation. URL-level churn is high and domain-level presence is stickier, which points the work at entity architecture rather than at individual pages. See Entity & Knowledge Architecture.