Executive summary

Google's public documentation on AI Overviews and AI Mode is more specific than most SEOs give it credit for. It names a real technique — "query fan-out" — defines it precisely, and gives worked examples. It states plainly that these features are "rooted in our core Search ranking and quality systems," meaning there is no secret parallel index. It has published, in writing, that llms.txt is ignored entirely: "Google Search ignores them," neither helping nor hurting rankings (Confidence: High — verbatim from Google Search Central's AI-optimization guide.). On the retrieval architecture, Google has been unusually forthcoming.

On one specific question, it has said nothing at all: once fan-out retrieves a set of candidate pages from the index, how does the synthesis layer decide which ones get quoted, linked, and surfaced — and which of the other equally-indexed, equally-eligible pages don't? Every official statement on this stops at "rooted in our core ranking systems" and goes no further. That's not an oversight after two years of AI Overviews in production. It reads as a deliberate line Google has chosen not to cross, for reasons that are defensible (publishing a scoring formula invites gaming) but that leave the industry's most consequential question — why does a page get cited — officially unanswered.

Two things compound the gap. First, Google spent from at least October 2024 telling SEOs there was "no plan" to break out AI Overview performance data in Search Console — a stance it reversed by mid-2026, shipping impressions-only reporting once pressure had built for long enough. Second, a meaningful share of the "Google AI Overview patents" circulating in SEO commentary turn out, on direct inspection of the patent record, to belong to other companies entirely — Glean and Microsoft, not Google. Real transparency on plumbing, engineered silence on scoring, and a citation ecosystem that hasn't checked its own homework: that's the actual state of what's known.

What Google has disclosed, and when

DateDisclosureSource
Oct 2024Danny Sullivan, on Search Console AI Overview data: "Can we get stats from things in Search Console from AI Overviews? No, and there's no plan to do that … The company views these all as part of search and not something that's worth breaking out."Aleyda Solis Q&A
Mar 2025AI Mode launches.Google
May 20, 2025Elizabeth Reid (VP, Head of Search): "Under the hood, AI Mode uses our query fan-out technique, breaking down your question into subtopics and issuing a multitude of queries simultaneously on your behalf." Deep Search escalates to "hundreds of searches."Google
Sept 26, 2023 (granted)US11769017B1, "Generative summaries for search results" — a real, granted Google LLC patent describing source selection, summary generation, and confidence-annotated linkification.Google Patents
Dec 3, 2024 (granted)US12158907B1, "Thematic search" — a real, granted Google LLC patent describing theme generation and hierarchical topic drill-down.Google Patents
Jan 2026Danny Sullivan warns against fragmenting content into "bite-sized chunks" for LLMs: "we don't want you to do that."Reported via industry press
Apr 21, 2026Search Central Live Toronto (the series' first Canadian stop): Google staff confirm blocking the Google-Extended crawler does not prevent a page appearing in AI Overviews/AI Mode, because pages already in the core index are pulled into fan-out regardless; data-nosnippet described as the only granular exclusion control.Conference recap, named Google presenters (Danny Sullivan, Martin Splitt, Daniel Waisberg)
May 15, 2026Google Search Central publishes its first vendor-neutral "AI optimization guide," authored under John Mueller's byline — the clearest single official statement of what does and doesn't affect AI features.Google Search Central
Jun 2026Google ships Gen-AI performance reports in Search Console — impressions for AI-surface appearances, but not clicks — reversing the 2024 "no plan" position.Google Search Central Blog

Part 1 — What Google will tell you

Read the actual documentation rather than the SEO commentary about it, and Google has disclosed more of the pipeline than it's usually credited for. Query fan-out is named and defined in plain language: "a set of concurrent, related queries generated by the model to request more information and fetch additional relevant search results to address the user's query" (Confidence: High — verbatim, Google Search Central.). The worked example Google gives is mundane on purpose — "how to fix a lawn full of weeds" spins out into "best herbicides for lawns," "remove weeds without chemicals," and similar sub-queries — precisely because the mechanism is meant to look procedural, not mysterious.

Google also states, without hedging, that generative features "rely on our core Search ranking systems to retrieve relevant, up-to-date web pages from our Search index" and are "rooted in our core Search ranking and quality systems" (Confidence: High — verbatim.). Practically, that rules out a specific and common misconception: there is no separate, secret "AI index" with its own admission criteria. Eligibility to appear as a supporting link requires the same baseline as classic search — the page must be indexed, eligible to show a snippet, and meet standard technical requirements — plus the same pre-existing controls site owners already had (nosnippet, data-nosnippet, max-snippet, noindex, robots.txt). No new AI-specific meta tag exists (Confidence: High — per the May 2026 AI-optimization guide and Search Console documentation.). And Google is explicit that llms.txt does nothing either way: "Google Search ignores them" (Confidence: High — verbatim.).

Part 2 — What Google will not tell you, and the reversal that proves it's a choice

Here is where the documentation stops, consistently, across every version Google has published: it never explains how the synthesis layer chooses among the candidates fan-out retrieves. If ten pages are equally indexed, equally snippet-eligible, and equally free of technical issues, nothing in Google's public materials describes what separates the one that gets quoted from the nine that don't. Every official answer redirects to "our core Search ranking and quality systems" — a phrase that describes an input, not a decision (Confidence: High — this is an absence, verified by the consistent absence of any further detail across every primary source checked: the 2025 blog posts, the January 2026 Search Central Live talks, and the May 2026 optimization guide.).

The clearest evidence that this silence is a policy rather than an accident is that Google has reversed course on adjacent transparency before, under sustained pressure, on its own timeline. In October 2024, Danny Sullivan told SEO consultant Aleyda Solis directly that Search Console would not break out AI Overview performance data: "No, and there's no plan to do that… The company views these all as part of search and not something that's worth breaking out." (Confidence: High — quote corroborated by two independent citations of the same interview; the primary recording itself was not directly re-accessed this session.) By June 2026, Google shipped exactly that — Gen-AI performance reports in Search Console, showing impressions for AI-surface appearances. Notably, impressions only, not clicks: the one number that would let a site quantify how much traffic AI Overviews are actually costing it remains absent even in the reversal. Google moved from "we won't break this out" to "we'll break out visibility, not outcome" — a real concession, calibrated to give away the least information that would still count as one.

What this pattern predicts: if the last two years are a guide, expect Google to eventually disclose more about post-fan-out selection only once external pressure — regulatory, competitive, or reputational — makes the current silence more costly than the disclosure would be. Nothing here suggests voluntary transparency is coming on its own schedule.

Part 3 — The patents that are and aren't Google's

Two real, granted Google patents do bear directly on this system, and are worth reading over any secondhand summary of them. US11769017B1, "Generative summaries for search results" (filed March 2023, granted September 2023), describes selecting source documents via a mix of query-dependent, query-independent, and user-dependent measures; generating a summary; linkifying portions of that summary back to source documents with annotated confidence levels; and routing between different generative models depending on query type to manage compute cost (Confidence: High — patent text directly reviewed via Google Patents.). US12158907B1, "Thematic search" (filed May 2023, granted December 2024), describes generating descriptive "themes" from a set of responsive documents and letting a user drill into hierarchical sub-themes, organizing results by topic rather than pure relevance rank (Confidence: High — same source.). Neither patent discloses a specific scoring formula for source selection; both describe architecture, not weights.

What's more instructive is what turns out not to be Google's. Two patents that circulate widely in SEO commentary as "how Google AI Overviews work" are misattributed. US20240256582A1, "Search with Generative Artificial Intelligence," is assigned to Glean Technologies Inc, an enterprise search company — not Google (Confidence: High — assignee field confirmed directly on the patent record.). WO2025128245A1, "Generative search engine results documents," is assigned to Microsoft Technology Licensing LLC — almost certainly describing Bing/Copilot's architecture, not Google's, despite surfacing prominently in searches for "Google AI Overview patent" (Confidence: High — same verification method.). A body of published SEO advice has been built, in part, on citing the wrong company's disclosure. That's a small, checkable fact with an outsized implication: read the assignee field before treating a patent as evidence of what Google does.

Part 4 — The one disclosure that isn't about clicks or chunking: Google-Extended doesn't do what site owners assume

Search Central Live's April 21, 2026 Toronto stop (covered in more depth elsewhere on this site for its Sullivan/chunking angle) produced a second, narrower disclosure that's easy to miss: Google staff — including Danny Sullivan and Martin Splitt — confirmed that blocking the Google-Extended crawler does not remove a page from AI Overviews or AI Mode results, because pages already present in the core Search index are pulled into fan-out retrieval regardless of the Google-Extended directive (Confidence: Medium — sourced via a conference attendee's slide recap rather than an official Google transcript, but consistent with Google's own written statement elsewhere that these features draw on the standard index.). The same session named data-nosnippet as effectively the only granular, page-level control for keeping specific content out of AI-generated responses — a narrower tool than many site owners assume they have when they set a Google-Extended directive and consider the matter closed. No timeline was given for further Search Console reporting beyond the impressions data that shipped two months later.

Counter-argument, taken seriously

Withholding the scoring mechanism is defensible, not just self-serving. Google has never published the classic web-ranking algorithm's exact weights either, for the same reason it presumably won't publish this one: any sufficiently specific description of a scoring function becomes a target-practice manual for manipulation, and the two-decade history of link-scheme and content-farm gaming is the reason ranking systems stay partially opaque by design. Judged against that baseline, refusing to detail AI Overview source-selection isn't a new kind of secrecy — it's the same posture Google has always taken, applied to a new surface. The stronger version of this critique is not "Google should publish its formula," but narrower: Google could disclose outcome data (which pages got cited and how often) without disclosing the mechanism, and the two-year gap before even impressions-only reporting shipped is the part that's hard to defend on manipulation-risk grounds alone, since impressions data carries none of the gaming risk that a scoring formula would.

What everyone is missing

Reading Google's silence as an absence of information, rather than a decision, misreads the pattern. The Search Console reversal shows Google will move — but only after sustained pressure, and only as far as the minimum concession that still counts as transparency. Treat every "no plan to do that" as a starting position, not a permanent one.

The patent-attribution sloppiness in GEO commentary is itself evidence of how young and undisciplined this specialty still is. If two of the most commonly cited "Google AI Overview patents" belong to Glean and Microsoft, a meaningful share of published advice built on "here's what Google's patent says" has been reasoning from the wrong company's document. This is not an isolated case, either — a separate audit of entity-resolution patents commonly attributed to Google's Knowledge Graph turns up the same error with a different set of filings; see Google Never Documented How AI Overviews Use Entities — So the Industry Cited Microsoft's Patents Instead. Verify the assignee field before treating any patent as proof of Google's mechanics — twice now, across two unrelated investigations, it hasn't been Google's.

The real technical requirement hasn't changed, and Google has said so directly. No new meta tag, no special schema, no llms.txt parsing, no separate index. The May 2026 optimization guide is Google's clearest statement yet that the fundamentals — indexability, snippet eligibility, ordinary technical health — are still the whole requirement for eligibility. What's undisclosed is what happens after eligibility, among the pages that already clear that bar.

Future predictions

  • Search Console's AI reporting adds more dimensions, slowly. Given the two-year gap between "no plan" and impressions-only reporting, expect click or conversion-adjacent data, if it comes, to arrive in a similarly delayed, minimal-viable form rather than all at once. (Confidence: Medium.)
  • Source-selection mechanics remain undisclosed through at least the next product cycle. Nothing in the current pattern suggests Google intends to publish scoring detail voluntarily. (Confidence: Medium-High.)
  • More misattributed "Google patents" will surface in SEO commentary before the practice self-corrects. As patent filings in this space accelerate across Google, Microsoft, OpenAI, Perplexity, and enterprise search vendors, the base rate for confused attribution rises, not falls. (Confidence: Medium.)
  • Google-Extended's limits become better understood and more frequently cited correctly. The Toronto disclosure — that Google-Extended doesn't exempt indexed pages from fan-out — is specific enough that expect it to circulate more widely as site owners test their own opt-out assumptions. (Confidence: Medium.)

Practical takeaways

  1. Stop treating llms.txt as a ranking lever. Google's own documentation states it is ignored entirely. Ship it only if it's free; never report it as a growth initiative.
  2. Verify patent assignees before citing "Google's patent" in a strategy deck. Two of the most-circulated ones aren't Google's. Check the assignee field on Google Patents directly.
  3. Don't assume Google-Extended blocks AI Overview inclusion. If a page is in the core index, it's eligible for fan-out retrieval regardless of that directive; data-nosnippet is the narrower, more reliable exclusion tool.
  4. Treat Search Console's Gen-AI impressions report as a visibility signal, not a performance report. It shows reach, not outcome — plan measurement accordingly, and see Zero-Click Is Old News. Zero-Visit Is the Real Crisis. for the traffic side of this same asymmetry.
  5. Invest in the fundamentals Google has explicitly confirmed still matter — indexability, snippet eligibility, technical health — over speculative "AI-specific" tactics. See Technical Foundations.
  6. Track how you're actually named and cited across engines directly, since Google won't tell you. See LLM Visibility Monitoring.