Executive summary
Everyone in SEO now has a favourite statistic about how to get cited by ChatGPT, Perplexity, or Google's AI Overviews, and the most repeated one — a "GEO can boost visibility by up to 40%" figure — comes from a single 2024 paper that ran its core benchmark against five sources pulled from Google and a GPT-3.5-turbo model standing in for a live AI engine, not against any production system in current use (Confidence: High — directly fetched arXiv 2311.09735, Section 3.1). That paper did also run a smaller live check against real Perplexity.ai in 2024, and it partially held up — quotation additions gained 22%, statistics gained up to roughly 37% — but keyword stuffing, which was merely neutral in the simulation, actually performed 10% worse live, and the whole check was a single snapshot on an architecture Perplexity has since rebuilt (Confidence: High — same paper, Section 6). That gap between what the famous study measured and what the industry now treats as settled instruction is the actual story.
Meanwhile the production data that has accumulated since — from Ahrefs, Otterly.ai, and independent chunking research — tells a more specific and less flattering story than "write authoritative content with statistics." Schema markup, the single most commonly prescribed GEO tactic, produced no meaningful citation uplift in a controlled difference-in-differences test of 1,885 real pages against 4,000 matched controls; AI Overview citations actually declined 4.6% after schema was added (Confidence: High — Ahrefs DiD study, 11 May 2026). Google's own documentation says as much directly: "There's also no special schema.org structured data that you need to add" (Confidence: High — Google Search Central, verbatim). What does correlate with citation, repeatedly and across independent datasets, is chunk-level structure and off-site brand presence — not document-level authority, and not markup.
The mechanics that actually decide which source gets quoted are retrieval mechanics: how a passage is chunked, whether it's the kind of passage an engine's fan-out queries would surface, and whether the brand already has a footprint on the platforms — YouTube, Reddit, LinkedIn — that these systems increasingly draw from directly (Confidence: High — Ahrefs most-cited domains tracker, 21 July 2026; Otterly 2026 citation report). None of this requires abandoning content quality. It does require abandoning the idea that a widely-quoted percentage from one 2024 paper is a specification sheet for 2026 engines.
A timeline / comparison table
| Date | Source | Sample | Key finding | Confidence |
|---|---|---|---|---|
| 28 Jun 2024 (v3) | GEO paper, KDD 2024 (arXiv 2311.09735) | 10,000 queries, 9 datasets/25 domains, simulated top-5 Google + GPT-3.5-turbo | "Up to 40%" visibility gain; range ~30–44% by tactic/metric | High |
| 2024 (Sec. 6) | Same paper, live validation | Real Perplexity.ai, one 2024 snapshot | Quotation Addition +22%, Statistics ~+37%; keyword stuffing −10% (diverged from simulation) | High |
| 3 Jul 2024 | Chroma Research | 328,208 tokens, 5 corpora, 472 queries | Recall varied 83.6–91.9% by chunking strategy alone (~8pt swing) | High |
| 26 May 2025 | Ahrefs brand-correlation study | 75,000 brands | Branded mentions correlate 0.664 with AI visibility vs. 0.218 for backlinks | High |
| 25 Jun 2025 | Seer Interactive | 5,000+ cited URLs via Peec.ai | Citation recency varies sharply by platform (Perplexity ~50% within a year) and vertical | Medium |
| 12 Dec 2025 | Ahrefs follow-up | Same 75,000-brand base, DR>40 filter | YouTube mentions (0.737) beat branded web mentions as top correlate | High |
| Mar 2026 | arXiv 2603.06976 | 36 chunking strategies, 6 domains, 5 embedding models | Paragraph Group Chunking: nDCG@5 ~45.9%, Hit@5 ~59% | High |
| 2026 | Ahrefs, AI Overviews vs. AI Mode | 540,000 query pairs | Only 13.7% citation overlap between the two systems despite 86–89.7% semantic agreement | High |
| 11 May 2026 | Ahrefs schema DiD study | 1,885 treated / 4,000 control pages | Schema addition: AIO −4.6% (significant), AI Mode +2.4% (n.s.), ChatGPT +2.2% (n.s.) | High |
| 2026 (Jan–Feb) | Otterly.ai citation report | 1M+ citations | Community platforms (Reddit, Quora) = 52.5% of all citations, ahead of brand domains | High |
| 3 Jun 2026 | Otterly.ai LinkedIn study | 1,310,455 citations, 6 platforms | LinkedIn content citations concentrated in 161,440 unique URLs | High |
| 21 Jul 2026 | Ahrefs most-cited domains tracker | 3M+ US queries | YouTube 21.1%, Reddit 18.5%, Facebook 10.7% lead all mention share | High |
Part 1 — The paper everyone quotes and almost no one reads past the abstract
The paper behind the "GEO can boost visibility by up to 40%" number — formally "GEO: Generative Engine Optimization," accepted at KDD 2024 — built something called GEO-bench: 10,000 queries pulled from nine source datasets (MS MARCO, Natural Questions, ELI-5, and others) spanning 25 topical domains, testing nine content tactics from "Cite Sources" to "Keyword Stuffing" (Confidence: High — arXiv 2311.09735, directly fetched). That is a serious, carefully constructed benchmark. It is also, by the authors' own description, a simulation: "only the top 5 sources are fetched from the Google search engine for every query," and the response synthesizing those five sources is generated by GPT-3.5-turbo standing in for the AI engine (Confidence: High — verbatim, Section 3.1). Five competing sources is a far smaller and more artificial pool than the hundreds or thousands of ranked pages a real search index draws from, and independent commentary has argued this inflates the relative-gain percentages the paper reports (Confidence: Medium — Blck Alpaca analysis).
The correction the industry conversation usually skips: this is not a paper that ignored live engines entirely. Section 6 ran a smaller validation against actual Perplexity.ai in 2024, and results partially replicated — Quotation Addition improved 22%, Cite Sources and Statistics Addition improved up to roughly 9% and 37% respectively (Confidence: High — Section 6, directly fetched). But one tactic inverted: keyword stuffing, neutral in simulation, measured 10% worse than baseline on live Perplexity (Confidence: High). That single divergence matters more than it looks — it's direct evidence that a tactic's simulated performance and its live performance are not the same number, on the one engine where both were actually tested. And that test happened once, on one platform, in 2024, on a Perplexity architecture that has changed substantially since. Treating the paper's headline percentages as current operating instructions for ChatGPT, Gemini, or 2026 AI Overviews — engines it never tested at all — is the error, not the paper's existence.
Diagram/table note for design: a two-column comparison graphic works well here — "Simulated GEO-bench (10,000 queries, GPT-3.5, 5 sources)" on the left against "Live Perplexity check (Section 6, 2024, single snapshot)" on the right, with the keyword-stuffing divergence flagged as the one row where the two columns disagree in direction, not just magnitude.
Part 2 — Schema markup: the most prescribed tactic with the least evidence behind it
If there is one instruction repeated across GEO advice more than any other, it's "add comprehensive schema markup." Google's own documentation contradicts the premise directly: "There are no additional technical requirements to appear in AI Overviews or AI Mode... There's also no special schema.org structured data that you need to add" (Confidence: High — Google Search Central, verbatim). Ahrefs tested the claim empirically rather than taking Google's word for it, tracking 1,885 pages that added JSON-LD schema between August 2025 and March 2026 against 4,000 matched control pages, all of which already had 100+ AI Overview citations before treatment — a difference-in-differences design built specifically to isolate schema's effect from general growth. The result: Google AI Overview citations declined 4.6%, a statistically significant drop; AI Mode moved +2.4% and ChatGPT +2.2%, neither significant. Their own conclusion: "Adding schema produced no major uplift in citations on any platform" (Confidence: High — Ahrefs, 11 May 2026, verbatim quote). That study has a real limit, too: it measured pages that were already being cited, not whether schema helps a previously invisible page get crawled and indexed in the first place.
A widely-repeated companion figure — that Article/FAQPage schema correlates with "2.3x more AI Overview citations," usually attributed to Semrush — could not be traced to any locatable primary Semrush report or methodology. Conflicting figures found elsewhere put comparable schema effects closer to 20–30%, not 2.3x (Confidence: Low — no verifiable primary source; flagged here specifically because it circulates as fact). Separately, a December 2024 study attributed to Search/Atlas reportedly found no correlation between schema coverage and citation rates at all, though the underlying study's sample size and methodology couldn't be independently verified beyond the secondary Search Engine Land writeup (Confidence: Low). The pattern across all three data points, strong and weak alike, is the same direction: nobody has produced solid evidence that schema markup drives citations, and the one rigorously controlled test found a negative effect on the platform that matters most.
Part 3 — Citations are won or lost at the chunk, not the document
The mechanism that actually seems to determine citation is retrieval at the passage level. Chroma Research tested chunking strategies across five corpora totalling 328,208 tokens and 472 queries, and found recall at five retrievals swung from 83.6% to 91.9% depending purely on how the source text was split — an eight-point spread with no change to the underlying content (Confidence: High — Chroma Research, 3 July 2024). Content-aware chunkers beat naive fixed-length splitting; notably, OpenAI's own default retrieval configuration (800 tokens, 400-token overlap) scored below average on recall and lowest on precision among every strategy tested (Confidence: High). A separate, more recent academic paper benchmarked 36 chunking strategies across six knowledge domains using five embedding models, with a large language model (chosen for strong correlation to human judges, Pearson r=0.879) grading relevance blind to the query text. Paragraph Group Chunking topped the field, with a Hit@5 of roughly 59% — meaning the correct passage appeared somewhere in the top five results about 59% of the time — against a normalized ranking quality score (nDCG@5) of roughly 45.9% for the same strategy (Confidence: High — arXiv 2603.06976, corrected: 59% is Hit@5, not nDCG@5, which is the lower figure).
SEO consultant Olaf Kopp, drawing on a 2024 Google patent on thematic search and a 2025 RAG paper, frames this as two separable factors that determine citation-worthiness: "LLM Readability," meaning how cleanly a passage can be extracted and structured, and "Chunk Relevance," meaning how well that specific passage matches the query — independent of whether the source document as a whole is authoritative. His formulation, verbatim: "Sources whose documents are less relevant than others can still be cited — if their chunks are better structured or more relevant" (Confidence: High — direct quote, verified). That single sentence is a more useful operating principle than any percentage from the GEO paper, because it explains something the schema data can't: why a page with thin domain authority but one perfectly self-contained, well-bounded paragraph can out-cite a comprehensive, schema-tagged competitor whose relevant sentence is buried mid-paragraph.
Counter-argument taken seriously: if chunk relevance dominates, why do brand-mention correlations (Part 4) matter at all — shouldn't a well-structured passage from an unknown site cite just as well as one from a known brand? The honest answer is that these are not competing explanations but sequential filters: retrieval-stage chunk relevance decides which candidate passages even reach the answer-generation step, while brand/entity signals appear to influence which of several retrieval-eligible candidates the model actually selects to quote or trust when multiple chunks are comparably relevant. The data supports both operating simultaneously; it does not support either one alone as sufficient.
Part 4 — Brand mentions and community platforms are outperforming both backlinks and schema
Ahrefs' first brand-correlation study, covering 75,000 brands, found unlinked branded web mentions correlate with AI Overview visibility at 0.664, compared with 0.218 for backlinks — roughly a 3:1 gap in favour of mentions (Confidence: High — Ahrefs, 26 May 2025). A December 2025 follow-up on the same brand base went further: YouTube mentions specifically were the single strongest correlate of AI visibility across ChatGPT, AI Mode, and AI Overviews, at 0.737 — beating even branded web mentions (Confidence: High — Ahrefs, 12 Dec 2025). That YouTube signal lines up with Ahrefs' separate most-cited-domains tracker, where YouTube alone accounts for 21.1% of citation mention share across 3M+ tracked queries, with Reddit at 18.5% and Facebook at 10.7% — three platforms accounting for roughly half of all citation mentions between them (Confidence: High — Ahrefs, 21 July 2026).
Otterly.ai's 2026 citation report, built on more than a million citations across ChatGPT, Perplexity, and AI Overviews, found community platforms like Reddit and Quora made up 52.5% of citations overall, versus 47.5% for brand domains — with brand-domain reliance varying sharply by engine: 59.8% of Google AI Overview citations go to brand domains, versus 44.7% for ChatGPT and just 28.9% for Perplexity, which leans hardest on community sources (Confidence: High — Otterly.ai, 2026). The same report found an estimated 73% of analyzed sites had technical barriers — robots.txt exclusions, CDN configuration, JavaScript rendering — actively blocking AI crawler access, meaning a meaningful share of the industry is invisible to these systems for reasons that have nothing to do with content quality (Confidence: High). Otterly's separate YouTube-specific study found something that should discourage chasing view counts as a proxy for citation-worthiness: popularity metrics — views, likes, subscriber counts — correlate with citation frequency at essentially zero (r ≈ -0.03), and long-form video accounts for 94% of cited YouTube content versus 5.7% for Shorts (Confidence: High).
Part 5 — Freshness and platform divergence are bigger variables than most GEO advice admits
Seer Interactive's analysis of over 5,000 cited URLs found citation recency — not crawl-log bot-visit frequency, which is a separate and often-conflated metric — varies by platform: roughly 50% of Perplexity citations went to content published within the prior year, versus 44% for AI Overviews and just 31% for ChatGPT (Confidence: Medium — Seer Interactive, 25 June 2025). Recency bias also varies sharply by vertical: financial services content shows an extreme preference for recent sources, while some energy and home-improvement content dating back to 2004 is still being crawled and cited (Confidence: Medium).
Perhaps the most consequential finding for anyone treating "AI search" as one target: Ahrefs found Google AI Overviews and Google AI Mode — two systems from the same company, running on the same query — share only 13.7% citation overlap (16.3% limited to top-3 citations), despite reaching semantically similar conclusions 86–89.7% of the time (Confidence: High — Ahrefs, 2026). In other words, the two systems usually agree on what to say but disagree on whom to credit for saying it. Google's own documentation offers a partial explanation, describing a "query fan-out" technique — "issuing multiple related searches across subtopics and data sources" — that is explicitly designed to "display a wider and more diverse set of helpful links... than with a classic web search" (Confidence: High — Google, verbatim). Ahrefs also found content length is close to irrelevant: across 560,346 AI Overviews and 1.68 million cited URLs, the correlation between word count and citation likelihood was essentially zero (Spearman r=0.04), with 53.4% of citations going to pages under 1,000 words (Confidence: High).
What everyone is missing
The industry keeps asking "what does the AI want" as if there is one AI. The 13.7% citation overlap between Google's own two AI systems is the clearest evidence yet that optimizing for "AI search" as a single target is a category error — a page can be highly citable to AI Mode's fan-out retrieval and functionally invisible to AI Overviews' narrower selection, using the same content, on the same day.
Chunk-level structure is being treated as a writing-style tip when the data says it's an information-architecture problem. Chroma's eight-point recall swing and the 36-strategy academic benchmark aren't about better prose — they're about where paragraph and section boundaries fall relative to where a query's fan-out subtopics land. That's an editorial and CMS-template decision, not a copywriting one, and almost none of the public GEO advice frames it that way.
Schema markup's popularity as a GEO tactic has outrun its evidence by a wide margin. Three independent data points — Google's own documentation, Ahrefs' controlled DiD study, and the unverifiable state of the "2.3x" Semrush figure — all point the same direction, yet schema remains the first line item on most GEO checklists, likely because it's the easiest tactic to sell as a deliverable, not because it's the one supported by data.
Brand-mention correlation is being read backwards. The Ahrefs findings don't mean "get mentioned more and you'll get cited more" in a simple causal sense — correlation this strong more plausibly reflects that brands already embedded in the community conversation (Reddit threads, YouTube explainers) are the ones whose content chunks are already being surfaced and validated by the same retrieval systems being measured. Mentions and citations may be two readings of the same underlying visibility, not cause and effect.
Future predictions
- Ahrefs, Otterly, or a comparable vendor will publish a chunk-level citation study within 12–18 months attempting to directly correlate CMS-level paragraph structure with citation rate, closing the gap between the academic chunking papers and production SEO data (Confidence: Medium — inferred from current research trajectory, not yet published).
- The 13.7% AI Overviews/AI Mode citation overlap will likely narrow somewhat as both systems mature on shared infrastructure, but a full convergence is unlikely given they appear to use genuinely different retrieval fan-out patterns (Confidence: Low — extrapolation, no forward-looking data available).
- Expect more corrections like Ahrefs' schema study to walk back earlier, less rigorous GEO claims as vendors move from correlational studies to controlled difference-in-differences designs (Confidence: Medium — based on the demonstrated trend from Ahrefs' own 2025 to 2026 studies).
- Community-platform reliance (Reddit, YouTube, LinkedIn) is more likely to increase than decrease across engines, given Otterly's finding that three-quarters of brand sites currently block AI crawlers entirely, leaving community-hosted mentions as the path of least resistance for these systems (Confidence: Medium).
- Academic hallucinated-citation research (146,932 estimated fabricated citations in 2025 papers alone) suggests citation verification tooling — for both academic and commercial contexts — will become a standalone product category rather than a footnote in retrieval research (Confidence: Medium — based on arXiv 2605.07723's scale, not a stated industry roadmap).
Practical takeaways
- Stop treating schema markup as a citation lever. Ahrefs' controlled test found a decline in AI Overview citations after adding it — spend that engineering time on Technical Foundations work that actually affects crawlability, like removing the robots.txt and JavaScript-rendering barriers Otterly found blocking 73% of sites.
- Audit content at the paragraph level, not the page level. Chroma's data shows an eight-point recall swing from chunking strategy alone — review whether your CMS templates produce self-contained, quotable paragraphs or long, context-dependent blocks, a core part of Entity & Knowledge Architecture.
- Build unlinked brand presence deliberately, especially on YouTube and Reddit — Ahrefs' 0.737 correlation for YouTube mentions and Otterly's 52.5% community-citation share both point the same direction, independent of backlink volume.
- Do not assume AI Overviews and AI Mode are the same target. With only 13.7% citation overlap, content strategy needs LLM Visibility Monitoring that tracks each engine separately rather than a single blended "AI visibility" score.
- Read the GEO paper's live-Perplexity section (Section 6) before quoting its headline percentages — the one tactic tested both ways (keyword stuffing) inverted between simulation and live results, which should discourage treating any single 2024 percentage as current instruction.
- Route citation strategy through Generative Engine Optimization work that treats chunk relevance and brand-entity signals as the two operative levers, rather than defaulting to schema and word count, which the data no longer supports.
