Skip to content
LumiRank
The journal
Technical13 min read

Google explained how AI Overviews retrieve content. It has never explained how they pick a winner.

Query fan-out is documented, defined, and demoed in Google's own developer docs. What happens after fan-out retrieves candidate pages has never been disclosed — and two "Google AI Overview patents" widely cited in SEO content belong to Glean and Microsoft.

By Dmytro Hrysiuk
A brass pneumatic-tube manifold on a plaster wall: one glass tube carrying a white envelope splits into about twenty tubes each carrying their own envelope, all converging into a sealed matte-black box that emits a single envelope on the far side

Executive summary

Google's public documentation on AI Overviews and AI Mode is more specific than most SEOs give it credit for. It names a real technique, "query fan-out," defines it precisely, and gives worked examples. It states plainly that these features are "rooted in our core Search ranking and quality systems," so there is no secret parallel index. It has said in writing that llms.txt is ignored entirely: "Google Search ignores them," neither helping nor hurting rankings. On the retrieval architecture, Google has been unusually forthcoming.

On one specific question it has said nothing at all. Once fan-out retrieves a set of candidate pages from the index, how does the synthesis layer decide which ones get quoted and linked, and which of the other equally-indexed, equally-eligible pages don't? Every official statement stops at "rooted in our core ranking systems" and goes no further. After two years of AI Overviews in production, that looks less like an oversight than a line Google has chosen not to cross. The reasons are defensible enough (publishing a scoring formula invites gaming), but they leave the industry's most consequential question, why a page gets cited, officially unanswered.

This piece sits inside our work on technical SEO, which is where the practice behind it is set out in full.

Two things compound the gap. First, Google spent from at least October 2024 telling SEOs there was "no plan" to break out AI Overview performance data in Search Console, a stance it reversed by mid-2026, shipping impressions-only reporting once pressure had built for long enough. Second, a meaningful share of the "Google AI Overview patents" circulating in SEO commentary turn out, on direct inspection of the patent record, to belong to other companies entirely: Glean and Microsoft, not Google. Google is transparent about the plumbing and silent about the scoring. And the citation ecosystem hasn't checked its own homework.

What Google has disclosed, and when

DateDisclosureSource
Oct 2024Danny Sullivan, on Search Console AI Overview data: "Can we get stats from things in Search Console from AI Overviews? No, and there's no plan to do that … The company views these all as part of search and not something that's worth breaking out."Aleyda Solis Q&A
Mar 2025AI Mode launches.Google
May 20, 2025Elizabeth Reid (VP, Head of Search): "Under the hood, AI Mode uses our query fan-out technique, breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf." Deep Search escalates to "hundreds of searches."Google
Sept 26, 2023 (granted)US11769017B1, "Generative summaries for search results" — a real, granted Google LLC patent describing source selection, summary generation, and confidence-annotated linkification.Google Patents
Dec 3, 2024 (granted)US12158907B1, "Thematic search" — a real, granted Google LLC patent describing theme generation and hierarchical topic drill-down.Google Patents
Jan 2026Danny Sullivan warns against fragmenting content into "bite-sized chunks" for LLMs: "we don't want you to do that."Reported via industry press
Apr 21, 2026Search Central Live Toronto (the series' first Canadian stop): Google staff confirm blocking the Google-Extended crawler does not prevent a page appearing in AI Overviews/AI Mode, because pages already in the core index are pulled into fan-out regardless; data-nosnippet described as the only granular exclusion control.Conference recap, named Google presenters (Danny Sullivan, Martin Splitt, Daniel Waisberg)
May 15, 2026Google Search Central publishes its first vendor-neutral "AI optimization guide," authored under John Mueller's byline — the clearest single official statement of what does and doesn't affect AI features.Google Search Central
Jun 2026Google ships Gen-AI performance reports in Search Console — impressions for AI-surface appearances, but not clicks — reversing the 2024 "no plan" position.Google Search Central Blog

Part 1 — What Google will tell you

Read the actual documentation rather than the SEO commentary about it, and Google has disclosed more of the pipeline than it usually gets credit for. Query fan-out is named and defined in plain language: "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query". The worked example is mundane on purpose: "how to fix a lawn full of weeds" spins out into "best herbicides for lawns," "remove weeds without chemicals," and similar sub-queries. Google wants the mechanism to look procedural, not mysterious.

Google also states, without hedging, that generative features "rely on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index" and are "rooted in our core Search ranking and quality systems". In practice that rules out a common misconception: there is no separate, secret "AI index" with its own admission criteria. To appear as a supporting link, a page must clear the same baseline as classic search. It has to be indexed, eligible to show a snippet, and meet standard technical requirements, and the only controls are the ones site owners already had (nosnippet, data-nosnippet, max-snippet, noindex, robots.txt). No new AI-specific meta tag exists. And Google is explicit that llms.txt does nothing either way: "Google Search ignores them".

Part 2 — What Google will not tell you, and the reversal that proves it's a choice

The documentation stops in the same place across every version Google has published: it never explains how the synthesis layer chooses among the candidates fan-out retrieves. If ten pages are equally indexed and equally snippet-eligible, nothing in Google's public materials describes what separates the one that gets quoted from the nine that don't. Every official answer redirects to "our core Search ranking and quality systems," a phrase that describes an input, not a decision.

The clearest evidence that this silence is policy rather than accident is that Google has reversed course on adjacent transparency before, under sustained pressure, on its own timeline. In October 2024, Danny Sullivan told SEO consultant Aleyda Solis directly that Search Console would not break out AI Overview performance data: "No, and there's no plan to do that… The company views these all as part of search and not something that's worth breaking out." By June 2026, Google shipped exactly that: Gen-AI performance reports in Search Console, showing impressions for AI-surface appearances. Impressions only, not clicks. The one number that would let a site quantify how much traffic AI Overviews are actually costing it remains absent even in the reversal. Google moved from "we won't break this out" to "we'll break out visibility, not outcome." A real concession, calibrated to give away the least information that would still count as one.

What this pattern predicts: if the last two years are a guide, expect Google to disclose more about post-fan-out selection only once external pressure, whether regulatory, competitive, or reputational, makes the current silence more costly than the disclosure would be. Nothing here suggests voluntary transparency is coming on its own schedule.

Part 3 — The patents that are and aren't Google's

Two real, granted Google patents bear directly on this system, and both are worth reading over any secondhand summary. US11769017B1, "Generative summaries for search results" (filed March 2023, granted September 2023), describes selecting source documents via a mix of query-dependent, query-independent, and user-dependent measures; generating a summary; linkifying portions of that summary back to source documents with annotated confidence levels; and routing between different generative models depending on query type to manage compute cost. US12158907B1, "Thematic search" (filed May 2023, granted December 2024), describes generating descriptive "themes" from a set of responsive documents and letting a user drill into hierarchical sub-themes, organizing results by topic rather than pure relevance rank. Neither patent discloses a specific scoring formula for source selection; both describe architecture, not weights.

More instructive is what turns out not to be Google's. Two patents that circulate widely in SEO commentary as "how Google AI Overviews work" are misattributed. US20240256582A1, "Search with Generative Artificial Intelligence," is assigned to Glean Technologies Inc, an enterprise search company. WO2025128245A1, "Generative search engine results documents," is assigned to Microsoft Technology Licensing LLC and almost certainly describes Bing/Copilot's architecture, not Google's, despite surfacing prominently in searches for "Google AI Overview patent". A body of published SEO advice has been built, in part, on citing the wrong company's disclosure. Read the assignee field before treating a patent as evidence of what Google does.

Part 4 — The one disclosure that isn't about clicks or chunking: Google-Extended doesn't do what site owners assume

Search Central Live's April 21, 2026 Toronto stop (covered in more depth elsewhere on this site for its Sullivan/chunking angle) produced a second, narrower disclosure that's easy to miss. Google staff, including Danny Sullivan and Martin Splitt, confirmed that blocking the Google-Extended crawler does not remove a page from AI Overviews or AI Mode results, because pages already present in the core Search index are pulled into fan-out retrieval regardless of the Google-Extended directive. The same session named data-nosnippet as effectively the only granular, page-level control for keeping specific content out of AI-generated responses. That is a narrower tool than many site owners assume they have when they set a Google-Extended directive and consider the matter closed. No timeline was given for further Search Console reporting beyond the impressions data that shipped two months later.

Counter-argument, taken seriously

Withholding the scoring mechanism is defensible, not just self-serving. Google has never published the classic web-ranking algorithm's exact weights either, for the same reason it presumably won't publish this one: any sufficiently specific description of a scoring function becomes a target-practice manual for manipulation, and two decades of link-scheme and content-farm gaming are why ranking systems stay partially opaque by design. Judged against that baseline, refusing to detail AI Overview source-selection is the same posture Google has always taken, applied to a new surface. The stronger version of this critique is narrower: Google could disclose outcome data (which pages got cited and how often) without disclosing the mechanism, and the two-year gap before even impressions-only reporting shipped is the part that's hard to defend on manipulation-risk grounds alone, since impressions data carries none of the gaming risk that a scoring formula would.

What everyone is missing

Reading Google's silence as an absence of information, rather than a decision, misreads the pattern. The Search Console reversal shows Google will move, but only after sustained pressure, and only as far as the minimum concession that still counts as transparency. Treat every "no plan to do that" as a starting position, not a permanent one.

The patent-attribution sloppiness in GEO commentary is itself evidence of how young and undisciplined this specialty still is. If two of the most commonly cited "Google AI Overview patents" belong to Glean and Microsoft, a meaningful share of published advice built on "here's what Google's patent says" has been reasoning from the wrong company's document. Nor is this an isolated case. A separate audit of entity-resolution patents commonly attributed to Google's Knowledge Graph turns up the same error with a different set of filings; see Google Never Documented How AI Overviews Use Entities — So the Industry Cited Microsoft's Patents Instead. Verify the assignee field before treating any patent as proof of Google's mechanics. Twice now, across two unrelated investigations, it hasn't been Google's.

The real technical requirement hasn't changed, and Google has said so directly. No new meta tag, no special schema, no llms.txt parsing, no separate index. The May 2026 optimization guide is Google's clearest statement yet that the fundamentals (indexability, snippet eligibility, ordinary technical health) are still the whole requirement for eligibility. What's undisclosed is what happens after eligibility, among the pages that already clear that bar.

Future predictions

  • Search Console's AI reporting adds more dimensions, slowly. Given the two-year gap between "no plan" and impressions-only reporting, expect click or conversion-adjacent data, if it comes, to arrive in a similarly delayed, minimal-viable form rather than all at once.
  • Source-selection mechanics remain undisclosed through at least the next product cycle. Nothing in the current pattern suggests Google intends to publish scoring detail voluntarily.
  • More misattributed "Google patents" will surface in SEO commentary before the practice self-corrects. As patent filings in this space accelerate across Google, Microsoft, OpenAI, Perplexity, and enterprise search vendors, the base rate for confused attribution rises, not falls.
  • Google-Extended's limits become better understood and more frequently cited correctly. The Toronto disclosure, that Google-Extended doesn't exempt indexed pages from fan-out, is specific enough to circulate widely as site owners test their own opt-out assumptions.

Practical takeaways

  1. Stop treating llms.txt as a ranking lever. Google's own documentation states it is ignored entirely. Ship it only if it's free; never report it as a growth initiative.
  2. Verify patent assignees before citing "Google's patent" in a strategy deck. Two of the most-circulated ones aren't Google's. Check the assignee field on Google Patents directly.
  3. Don't assume Google-Extended blocks AI Overview inclusion. If a page is in the core index, it's eligible for fan-out retrieval regardless of that directive; data-nosnippet is the narrower, more reliable exclusion tool.
  4. Treat Search Console's Gen-AI impressions report as a visibility signal, not a performance report. It shows reach, not outcome. Plan measurement accordingly, and see Zero-Click Is Old News. Zero-Visit Is the Real Crisis. for the traffic side of this same asymmetry.
  5. Invest in the fundamentals Google has explicitly confirmed still matter (indexability, snippet eligibility, technical health) over speculative "AI-specific" tactics. See Technical SEO.
  6. Track how you're actually named and cited across engines directly, since Google won't tell you. See LLM Visibility Monitoring.

Read the rest of the journal.

Key takeaways
  • Google's documentation clearly defines and demonstrates query fan-out, and confirms AI Overviews/AI Mode draw on the same core Search index and ranking systems — no secret parallel index exists.
  • Google has never disclosed how the synthesis layer scores or selects among the candidate pages fan-out retrieves; every official statement stops at "rooted in our core ranking systems."
  • Google reversed a 2024 "no plan" position on Search Console AI-performance data by shipping impressions-only reporting in mid-2026 — evidence the silence on measurement is a policy choice that moves under pressure, not a fixed limitation.
  • Two widely-cited "Google AI Overview patents" (US20240256582A1, WO2025128245A1) are actually assigned to Glean Technologies and Microsoft, not Google — a citation error worth checking before trusting any patent-based SEO claim.
  • Two patents genuinely belonging to Google (US11769017B1, US12158907B1) describe summary generation and thematic organization, but no scoring formula.
  • Google staff confirmed at Search Central Live Toronto (April 2026) that blocking Google-Extended does not exempt an indexed page from AI Overview/AI Mode fan-out retrieval.
  • `llms.txt` has no effect on Google rankings or AI features, per Google's own written documentation.

Frequently asked

Has Google disclosed how AI Overviews decide which pages to cite?
Partially. Google has clearly documented the retrieval mechanism — query fan-out — and confirmed that generative features draw on the same core Search index and ranking systems as classic search. It has never disclosed how the synthesis layer scores or selects among the candidate pages that fan-out retrieves. That specific step remains undocumented in every official source checked.
Is `llms.txt` a Google ranking factor?
No. Google's own Search Central documentation states plainly that Google Search ignores the file entirely — it "will neither harm nor help your site's visibility or rankings."
Are the "Google AI Overview patents" cited in SEO articles actually Google's?
Not always. Two commonly cited patents — US20240256582A1 and WO2025128245A1 — are assigned to Glean Technologies and Microsoft respectively, not Google. Two patents that are genuinely Google's and genuinely relevant are US11769017B1 ("Generative summaries for search results") and US12158907B1 ("Thematic search").
Does blocking Google-Extended keep my content out of AI Overviews?
No, according to Google staff at Search Central Live Toronto (April 2026). If a page is already present in Google's core Search index, it remains eligible for query fan-out retrieval regardless of a Google-Extended block. data-nosnippet is the narrower control that actually limits what appears in AI-generated summaries.
Does Search Console show AI Overview performance data now?
As of mid-2026, yes, partially — Google shipped Gen-AI performance reports showing impressions for AI-surface appearances. This reverses an October 2024 statement from Google's Search Liaison that there was "no plan" to break this data out. Click-level data for these surfaces is still not included.
From the journal

Most sites fail this at the crawl and rendering layer before content is ever the problem: