Skip to content
LumiRank
The journal
Research16 min read

The 40% citation boost in the famous GEO study was simulated. Here's what actually gets you cited.

The most-cited AI-citation study ran on GPT-3.5 and five fake search results. Ahrefs, Otterly, and Chroma's actual production data point somewhere else entirely.

By Dmytro Hrysiuk
Torn scraps of blank paper spread across a dark wooden desk under a warm lamp, with a hand reaching in to pick one up

Executive summary

Everyone in SEO has a favourite statistic about how to get cited by ChatGPT, Perplexity, or Google's AI Overviews. The most repeated one, the claim that "GEO can boost visibility by up to 40%," comes from a single 2024 paper. That paper ran its core benchmark against five sources pulled from Google, with GPT-3.5-turbo standing in for a live AI engine. No production system in current use was tested. The same paper did run a smaller live check against real Perplexity.ai in 2024, and it partially held up: quotation additions gained 22% and statistics gained up to roughly 37%. But keyword stuffing, merely neutral in the simulation, performed 10% worse live, and the whole check was a single snapshot on an architecture Perplexity has since rebuilt. What the famous study measured and what the industry now treats as settled instruction are two different things.

This piece sits inside our work on LLM visibility monitoring, which is where the practice behind it is set out in full.

The production data accumulated since then, from Ahrefs, Otterly.ai, and independent chunking research, tells a more specific and less flattering story than "write authoritative content with statistics." Schema markup, the single most commonly prescribed GEO tactic, produced no meaningful citation uplift in a controlled difference-in-differences test of 1,885 real pages against 4,000 matched controls. AI Overview citations declined 4.6% after schema was added. Google's own documentation says so directly: "There's also no special schema.org structured data that you need to add". What correlates with citation, repeatedly and across independent datasets, is chunk-level structure and off-site brand presence. Document-level authority and markup do not.

The mechanics that decide which source gets quoted are retrieval mechanics: how a passage is chunked, whether it is the kind of passage an engine's fan-out queries would surface, and whether the brand already has a footprint on the platforms these systems draw from directly (YouTube, Reddit, LinkedIn). None of this requires abandoning content quality. It does require dropping the idea that a widely quoted percentage from one 2024 paper is a specification sheet for 2026 engines.

A timeline / comparison table

DateSourceSampleKey findingConfidence
28 Jun 2024 (v3)GEO paper, KDD 2024 (arXiv 2311.09735)10,000 queries, 9 datasets/25 domains, simulated top-5 Google + GPT-3.5-turbo"Up to 40%" visibility gain; range ~30–44% by tactic/metricHigh
2024 (Sec. 6)Same paper, live validationReal Perplexity.ai, one 2024 snapshotQuotation Addition +22%, Statistics ~+37%; keyword stuffing −10% (diverged from simulation)High
3 Jul 2024Chroma Research328,208 tokens, 5 corpora, 472 queriesRecall varied 83.6–91.9% by chunking strategy alone (~8pt swing)High
26 May 2025Ahrefs brand-correlation study75,000 brandsBranded mentions correlate 0.664 with AI visibility vs. 0.218 for backlinksHigh
25 Jun 2025Seer Interactive5,000+ cited URLs via Peec.aiCitation recency varies sharply by platform (Perplexity ~50% within a year) and verticalMedium
12 Dec 2025Ahrefs follow-upSame 75,000-brand base, DR>40 filterYouTube mentions (0.737) beat branded web mentions as top correlateHigh
Mar 2026arXiv 2603.0697636 chunking strategies, 6 domains, 5 embedding modelsParagraph Group Chunking: nDCG@5 ~45.9%, Hit@5 ~59%High
2026Ahrefs, AI Overviews vs. AI Mode540,000 query pairsOnly 13.7% citation overlap between the two systems despite 86–89.7% semantic agreementHigh
11 May 2026Ahrefs schema DiD study1,885 treated / 4,000 control pagesSchema addition: AIO −4.6% (significant), AI Mode +2.4% (n.s.), ChatGPT +2.2% (n.s.)High
2026 (Jan–Feb)Otterly.ai citation report1M+ citationsCommunity platforms (Reddit, Quora) = 52.5% of all citations, ahead of brand domainsHigh
3 Jun 2026Otterly.ai LinkedIn study1,310,455 citations, 6 platformsLinkedIn content citations concentrated in 161,440 unique URLsHigh
21 Jul 2026Ahrefs most-cited domains tracker3M+ US queriesYouTube 21.1%, Reddit 18.5%, Facebook 10.7% lead all mention shareHigh

Part 1 — The paper everyone quotes and almost no one reads past the abstract

The paper behind the "GEO can boost visibility by up to 40%" number is "GEO: Generative Engine Optimization," accepted at KDD 2024. It built something called GEO-bench: 10,000 queries pulled from nine source datasets (MS MARCO, Natural Questions, ELI-5, and others) spanning 25 topical domains, with nine content tactics tested, from "Cite Sources" to "Keyword Stuffing". That is a serious, carefully built benchmark. It is also, by the authors' own description, a simulation: "only the top 5 sources are fetched from the Google search engine for every query," and GPT-3.5-turbo generates the response synthesizing those five sources, standing in for the AI engine. Five competing sources make a far smaller and more artificial pool than the hundreds or thousands of ranked pages a real search index draws from, and independent commentary has argued this inflates the relative-gain percentages the paper reports.

The correction the industry conversation usually skips: the paper did not ignore live engines entirely. Section 6 ran a smaller validation against actual Perplexity.ai in 2024, and the results partially replicated. Quotation Addition improved 22%; Cite Sources and Statistics Addition improved up to roughly 9% and 37% respectively. One tactic inverted, though. Keyword stuffing, neutral in simulation, measured 10% worse than baseline on live Perplexity. That divergence is direct evidence that a tactic's simulated performance and its live performance are not the same number, on the one engine where both were actually tested. And that test happened once, on one platform, in 2024, on a Perplexity architecture that has changed substantially since. The error is treating the paper's headline percentages as current operating instructions for ChatGPT, Gemini, or 2026 AI Overviews, engines it never tested at all. The paper itself is not the problem.

Diagram/table note for design: a two-column comparison graphic works well here: "Simulated GEO-bench (10,000 queries, GPT-3.5, 5 sources)" on the left against "Live Perplexity check (Section 6, 2024, single snapshot)" on the right, with the keyword-stuffing divergence flagged as the one row where the two columns disagree in direction, not just magnitude.

Part 2 — Schema markup: the most prescribed tactic with the least evidence behind it

The most repeated instruction in GEO advice is to add schema markup. Google's own documentation contradicts the premise directly: "There are no additional technical requirements to appear in AI Overviews or AI Mode... There's also no special schema.org structured data that you need to add". Ahrefs tested the claim empirically rather than taking Google's word for it. The study tracked 1,885 pages that added JSON-LD schema between August 2025 and March 2026 against 4,000 matched control pages, all of which already had 100+ AI Overview citations before treatment. That is a difference-in-differences design built specifically to isolate schema's effect from general growth. Google AI Overview citations declined 4.6%, a statistically significant drop. AI Mode moved +2.4% and ChatGPT +2.2%, neither significant. Their own conclusion: "Adding schema produced no major uplift in citations on any platform". The study has a real limit, too. It measured pages that were already being cited, not whether schema helps a previously invisible page get crawled and indexed in the first place.

A widely repeated companion figure, the claim that Article/FAQPage schema correlates with "2.3x more AI Overview citations," usually attributed to Semrush, could not be traced to any locatable primary Semrush report or methodology. Conflicting figures found elsewhere put comparable schema effects closer to 20–30%, not 2.3x. Separately, a December 2024 study attributed to Search/Atlas reportedly found no correlation between schema coverage and citation rates at all, though the underlying sample size and methodology couldn't be independently verified beyond the secondary Search Engine Land writeup. All three data points, strong and weak alike, point the same direction. Nobody has produced solid evidence that schema markup drives citations, and the one rigorously controlled test found a negative effect on the platform that matters most.

Part 3 — Citations are won or lost at the chunk, not the document

The mechanism that seems to determine citation is retrieval at the passage level. Chroma Research tested chunking strategies across five corpora totalling 328,208 tokens and 472 queries. Recall at five retrievals swung from 83.6% to 91.9% depending purely on how the source text was split, an eight-point spread with no change to the underlying content. Content-aware chunkers beat naive fixed-length splitting. OpenAI's own default retrieval configuration (800 tokens, 400-token overlap) scored below average on recall and lowest on precision among every strategy tested. A separate, more recent academic paper benchmarked 36 chunking strategies across six knowledge domains using five embedding models, with a large language model (chosen for strong correlation to human judges, Pearson r=0.879) grading relevance blind to the query text. Paragraph Group Chunking topped the field. Its Hit@5 was roughly 59%, meaning the correct passage appeared somewhere in the top five results about 59% of the time, against a normalized ranking quality score (nDCG@5) of roughly 45.9% for the same strategy.

SEO consultant Olaf Kopp, drawing on a 2024 Google patent on thematic search and a 2025 RAG paper, frames this as two separable factors that decide citation-worthiness. The first is "LLM Readability," how cleanly a passage can be extracted and structured. The second is "Chunk Relevance," how well that specific passage matches the query, independent of whether the source document as a whole is authoritative. His formulation, verbatim: "Sources whose documents are less relevant than others can still be cited — if their chunks are better structured or more relevant". That one sentence is a more useful operating principle than any percentage from the GEO paper, because it explains something the schema data can't: why a page with thin domain authority but one perfectly self-contained, well-bounded paragraph can out-cite a heavily schema-tagged competitor whose relevant sentence sits buried mid-paragraph.

Counter-argument taken seriously: if chunk relevance dominates, why do brand-mention correlations (Part 4) matter at all? Shouldn't a well-structured passage from an unknown site cite just as well as one from a known brand? The honest answer is that these are sequential filters rather than competing explanations. Retrieval-stage chunk relevance decides which candidate passages reach the answer-generation step at all. Brand and entity signals then appear to influence which of several retrieval-eligible candidates the model actually selects to quote or trust when multiple chunks are comparably relevant. The data supports both operating simultaneously; it does not support either one alone as sufficient.

Part 4 — Brand mentions and community platforms are outperforming both backlinks and schema

Ahrefs' first brand-correlation study, covering 75,000 brands, found unlinked branded web mentions correlate with AI Overview visibility at 0.664, compared with 0.218 for backlinks. That is roughly a 3:1 gap in favour of mentions. A December 2025 follow-up on the same brand base went further. YouTube mentions specifically were the single strongest correlate of AI visibility across ChatGPT, AI Mode, and AI Overviews at 0.737, beating even branded web mentions. That YouTube signal lines up with Ahrefs' separate most-cited-domains tracker, where YouTube alone accounts for 21.1% of citation mention share across 3M+ tracked queries, with Reddit at 18.5% and Facebook at 10.7%. Three platforms, roughly half of all citation mentions between them.

Otterly.ai's 2026 citation report, built on more than a million citations across ChatGPT, Perplexity, and AI Overviews, found community platforms like Reddit and Quora made up 52.5% of citations overall, versus 47.5% for brand domains. Brand-domain reliance varies sharply by engine: 59.8% of Google AI Overview citations go to brand domains, versus 44.7% for ChatGPT and just 28.9% for Perplexity, which leans hardest on community sources. The same report found an estimated 73% of analyzed sites had technical barriers (robots.txt exclusions, CDN configuration, JavaScript rendering) actively blocking AI crawler access. A meaningful share of the industry is invisible to these systems for reasons that have nothing to do with content quality. Otterly's separate YouTube-specific study should discourage anyone chasing view counts as a proxy for citation-worthiness. Popularity metrics (views, likes, subscriber counts) correlate with citation frequency at essentially zero, r ≈ -0.03, and long-form video accounts for 94% of cited YouTube content versus 5.7% for Shorts.

Part 5 — Freshness and platform divergence are bigger variables than most GEO advice admits

Seer Interactive's analysis of over 5,000 cited URLs found citation recency varies by platform. (Recency here means the age of the cited content, not crawl-log bot-visit frequency, a separate and often-conflated metric.) Roughly 50% of Perplexity citations went to content published within the prior year, versus 44% for AI Overviews and just 31% for ChatGPT. Recency bias also varies sharply by vertical. Financial services content shows an extreme preference for recent sources, while some energy and home-improvement content dating back to 2004 is still being crawled and cited.

Perhaps the most consequential finding for anyone treating "AI search" as one target: Ahrefs found Google AI Overviews and Google AI Mode, two systems from the same company running on the same query, share only 13.7% citation overlap (16.3% limited to top-3 citations), despite reaching semantically similar conclusions 86–89.7% of the time. The two systems usually agree on what to say. They disagree on whom to credit for saying it. Google's own documentation offers a partial explanation, describing a "query fan-out" technique, "issuing multiple related searches across subtopics and data sources," explicitly designed to "display a wider and more diverse set of helpful links... than with a classic web search". Ahrefs also found content length is close to irrelevant: across 560,346 AI Overviews and 1.68 million cited URLs, the correlation between word count and citation likelihood was essentially zero (Spearman r=0.04), with 53.4% of citations going to pages under 1,000 words.

What everyone is missing

The industry keeps asking "what does the AI want" as if there is one AI. The 13.7% citation overlap between Google's own two AI systems is the clearest evidence yet that optimizing for "AI search" as a single target is a category error. A page can be highly citable to AI Mode's fan-out retrieval and functionally invisible to AI Overviews' narrower selection, using the same content, on the same day.

Chunk-level structure is being treated as a writing-style tip when the data says it's an information-architecture problem. Chroma's eight-point recall swing and the 36-strategy academic benchmark have nothing to do with better prose. They concern where paragraph and section boundaries fall relative to where a query's fan-out subtopics land. That is an editorial and CMS-template decision, not a copywriting one, and almost none of the public GEO advice frames it that way.

Schema markup's popularity as a GEO tactic has outrun its evidence by a wide margin. Three independent data points (Google's own documentation, Ahrefs' controlled DiD study, and the unverifiable state of the "2.3x" Semrush figure) all point the same direction. Yet schema remains the first line item on most GEO checklists, likely because it is the easiest tactic to sell as a deliverable rather than the one supported by data.

Brand-mention correlation is being read backwards. The Ahrefs findings do not mean "get mentioned more and you'll get cited more" in any simple causal sense. Correlation this strong more plausibly reflects that brands already embedded in the community conversation (Reddit threads, YouTube explainers) are the ones whose content chunks are already being surfaced and validated by the same retrieval systems being measured. Mentions and citations may be two readings of the same underlying visibility rather than cause and effect.

Future predictions

  • Ahrefs, Otterly, or a comparable vendor will publish a chunk-level citation study within 12–18 months that attempts to directly correlate CMS-level paragraph structure with citation rate, closing the gap between the academic chunking papers and production SEO data.
  • The 13.7% AI Overviews/AI Mode citation overlap will likely narrow somewhat as both systems mature on shared infrastructure, but a full convergence is unlikely given they appear to use genuinely different retrieval fan-out patterns.
  • Expect more corrections like Ahrefs' schema study to walk back earlier, less rigorous GEO claims as vendors move from correlational studies to controlled difference-in-differences designs.
  • Community-platform reliance (Reddit, YouTube, LinkedIn) is more likely to increase than decrease across engines, given Otterly's finding that three-quarters of brand sites currently block AI crawlers entirely, leaving community-hosted mentions as the path of least resistance for these systems.
  • Academic hallucinated-citation research (146,932 estimated fabricated citations in 2025 papers alone) suggests citation verification tooling, for both academic and commercial contexts, will become a standalone product category rather than a footnote in retrieval research.

Practical takeaways

  1. Stop treating schema markup as a citation lever. Ahrefs' controlled test found a decline in AI Overview citations after adding it. Spend that engineering time on Technical SEO work that actually affects crawlability, like removing the robots.txt and JavaScript-rendering barriers Otterly found blocking 73% of sites.
  2. Audit content at the paragraph level, not the page level. Chroma's data shows an eight-point recall swing from chunking strategy alone. Review whether your CMS templates produce self-contained, quotable paragraphs or long, context-dependent blocks, a core part of Entity & Knowledge Architecture.
  3. Build unlinked brand presence deliberately, especially on YouTube and Reddit. Ahrefs' 0.737 correlation for YouTube mentions and Otterly's 52.5% community-citation share point the same direction, independent of backlink volume.
  4. Do not assume AI Overviews and AI Mode are the same target. With only 13.7% citation overlap, content strategy needs LLM Visibility Monitoring that tracks each engine separately rather than a single blended "AI visibility" score.
  5. Read the GEO paper's live-Perplexity section (Section 6) before quoting its headline percentages. The one tactic tested both ways, keyword stuffing, inverted between simulation and live results. That alone should discourage treating any single 2024 percentage as current instruction.
  6. Route citation strategy through Generative Engine Optimization work that treats chunk relevance and brand-entity signals as the two operative levers, rather than defaulting to schema and word count, which the data no longer supports.

Read the rest of the journal.

Key takeaways
  • Ahrefs' controlled study of 1,885 pages found adding schema markup produced a statistically significant 4.6% decline in Google AI Overview citations, and no significant effect on AI Mode or ChatGPT.
  • Google's own documentation states no special schema.org markup is required to appear in AI Overviews or AI Mode.
  • The widely-cited "GEO boosts visibility up to 40%" figure comes from a simulated benchmark (five Google sources, GPT-3.5-turbo); a smaller live 2024 Perplexity check partially replicated it but one tactic inverted direction.
  • Chunking strategy alone produced an 8.3-percentage-point swing in retrieval recall across Chroma Research's five-corpus test, with no change to the underlying content.
  • Ahrefs found branded mentions correlate with AI visibility at roughly 3x the strength of backlinks (0.664 vs. 0.218), and YouTube mentions specifically correlate even more strongly (0.737).
  • Google AI Overviews and Google AI Mode share only 13.7% citation overlap on identical queries despite 86–89.7% semantic agreement in their answers.
  • An estimated 73% of sites analyzed by Otterly.ai have technical barriers (robots.txt, CDN, JavaScript rendering) actively blocking AI crawler access.

Frequently asked

Does adding schema markup help get a page cited by ChatGPT or Google AI Overviews?
The best controlled evidence says no. Ahrefs tracked 1,885 pages that added JSON-LD schema against 4,000 matched control pages and found no significant uplift on ChatGPT or AI Mode, and a significant 4.6% decline on Google AI Overviews specifically. Google's own documentation states no special schema is required to appear in either feature.
Is the "GEO can boost visibility by 40%" statistic reliable?
It's real, but it comes from a simulated benchmark — five sources pulled from Google, synthesized by GPT-3.5-turbo, not a live production AI engine. The same paper did run a smaller live check on Perplexity.ai in 2024 that partially replicated the effect, but one tactic (keyword stuffing) performed in the opposite direction live versus simulated, and the check has not been repeated on today's engines.
What actually correlates with getting cited, if not schema or content length?
Chunk-level structure (how cleanly a passage can be extracted and matched to a query) and off-site brand presence, particularly on YouTube and Reddit, show the strongest and most consistently replicated correlations across independent studies from Ahrefs, Otterly.ai, and Chroma Research.
Do Google AI Overviews and Google AI Mode cite the same sources?
No. Ahrefs found only 13.7% citation overlap between the two systems on identical queries, despite reaching semantically similar answers 86–89.7% of the time — meaning the same content can be cited by one and ignored by the other.
Does longer, more comprehensive content get cited more often?
Ahrefs found essentially no correlation between word count and AI Overview citation likelihood (Spearman r=0.04), and 53.4% of all citations went to pages under 1,000 words.
How much does video/YouTube popularity matter for AI citations?
Very little on its own. Otterly.ai found near-zero correlation (r≈-0.03) between a video's views, likes, or subscriber count and how often it gets cited by AI systems — long-form video format matters far more than popularity metrics.
From the journal

Findings like this are only useful against your own numbers. That starts with a baseline: