Executive summary
Enterprise retrieval-augmented generation is no longer a pilot-stage technology. Microsoft disclosed more than 20 million paid Microsoft 365 Copilot seats as of its FY26 Q3 earnings call on April 29, 2026, up from 15 million just three months earlier — the fastest seat growth Copilot has seen since launch (Confidence: High — corroborated across Microsoft investor relations, TechCrunch, and NoJitter). Satya Nadella used that call to name Bayer, Johnson & Johnson, Mercedes and Roche as each running more than 90,000 Copilot seats, and announced an Accenture deal covering over 740,000 seats, calling it "our largest Copilot win to date" (Confidence: High — direct quote confirmed via TechCrunch). Glean, the enterprise search and agent platform competing directly with Copilot, raised a $150 million Series F at a $7.2 billion valuation in June 2025, with Wellington Management, Khosla Ventures and existing backers Sequoia and Lightspeed participating (Confidence: High — confirmed via Glean's own press release and independent corroboration).
That growth is forcing a question the SEO and content-marketing industry has mostly answered by assertion rather than evidence: what actually makes content retrievable and citable by a RAG system, whether that system is ChatGPT answering a consumer, or Copilot answering an employee inside SharePoint? A March 2026 controlled study from WordLift researchers, run across 2,443 query evaluations using Vertex AI Vector Search and Gemini models, found that adding JSON-LD/Schema.org markup to a page produced a statistically real but practically negligible accuracy gain — Cohen's d of 0.18, in the "small" range bordering on trivial (Confidence: High — primary source, arXiv:2603.10700v1). Rewriting the same underlying facts as legible, entity-dense prose instead of hidden markup produced a 29.6–29.8% accuracy gain in the same test conditions (Confidence: High — same primary source).
That gap is the story. The industry-standard claim that "schema markup gets you cited 2.5x more by AI systems" turns out to be unsourced — it circulates across at least seven SEO blogs with no primary study behind it, and the sites aren't even consistent with each other, some citing 2.5x and others 3.2x for the identical claim (Confidence: High — confirmed via direct multi-site search). Meanwhile actual enterprise RAG systems — the ones ingesting real B2B content at scale, right now, for money — are proving structurally indifferent to the tactic B2B marketers have spent two years being told to prioritize.
A timeline / comparison table
| Date | Event | Detail | Confidence |
|---|---|---|---|
| 2025-06-10 | Glean closes Series F | $150M raised at $7.2B valuation, led by Wellington Management | High — official press release, corroborated |
| 2026-01-28 | Microsoft FY26 Q2 earnings | 15M paid Microsoft 365 Copilot seats disclosed; GitHub Copilot reaches 4.7M paid subscribers (~+75% YoY) | High — Microsoft IR, TechCrunch, NoJitter |
| 2026-03-09 | KEO Marketing publishes B2B traffic-decline analysis | Claims 73% of B2B sites saw significant organic traffic loss 2024–2025, avg. 34% YoY decline; methodology undisclosed | Medium — traced to primary source, but it is the company's own unaudited marketing claim |
| 2026-03-11 | WordLift arXiv paper published | 2,443 evaluations across 4 domains; JSON-LD alone: Δ=+0.17, d=0.18; legible entity pages: +29.6–29.8% | High — primary source, fetched and verified verbatim |
| 2026-04-29 | Microsoft FY26 Q3 earnings | Paid Copilot seats pass 20M; Nadella names Bayer, J&J, Mercedes, Roche (90k+ seats each) and a 740,000-seat Accenture deal | High — Microsoft IR, TechCrunch |
| 2026-05-05 | Microsoft 2026 Work Trend Index Annual Report | Organizational rollout factors account for 67% of AI-driven productivity impact vs. 32% for individual usage factors | High — primary Microsoft WorkLab page |
| 2026 (undated) | SE Ranking / OrganiKPI llms.txt study | 10.13% adoption across ~300,000 domains studied; removing llms.txt as a model variable improved citation-prediction accuracy | Medium — primary page fetched twice, figures confirmed |
Part 1 — The enterprise RAG buildout is no longer speculative
Microsoft's own documentation describes Copilot Studio's retrieval pipeline in specific, mechanical terms: a query-rewriting step that folds in the last ten conversational turns, retrieval across knowledge sources — public web/Bing, SharePoint, OneDrive, uploaded files, Dataverse tables, Graph connectors, real-time connectors, and Azure AI Search — capped at the top three results per source, with security trimming applied on any source using delegated authentication, followed by summarization with citations and a moderation/grounding pass (Confidence: High — Microsoft Learn primary documentation). Short-term conversational state is retained for less than 30 days and explicitly is not used to train the underlying models (Confidence: High — same source). This is not a marketing description; it is the operational architecture running inside every enterprise that has turned Copilot on.
The adoption numbers back the architecture. Fifteen million paid Microsoft 365 Copilot seats in January 2026 became more than 20 million by the end of April — three months, five million net new paid seats, the fastest stretch of growth since Copilot's release (Confidence: High). GitHub Copilot, a separate product line disclosed on the earlier January call, hit 4.7 million paid subscribers, up roughly 75% year over year (Confidence: High). Glean, meanwhile, is scaling on the enterprise-search side of the same trend: a $7.2 billion valuation as of June 2025, and an independent estimate from research firm Sacra putting its annualized revenue near $300 million as of roughly May 2026 (Confidence: Medium — single secondary-source estimate, not company-disclosed).
Glean's own customer stories, self-published and therefore promotional rather than audited, still show a mechanism worth taking seriously: Duolingo employees built more than 500 internal AI agents within six weeks of launch (3,400+ total), reporting 500+ hours saved monthly and a claimed 5x ROI (Confidence: Medium — vendor-published, figures confirmed verbatim on the page but self-reported). One Duolingo engineer's quote is worth reading in full because it names the exact behavioural shift enterprise RAG is selling:
"Glean Chat is the most underrated feature within Glean. Similar to how some people now reach for ChatGPT before Google, Glean Chat can answer some questions even more effectively than a search can." — Art Chaidarun, Principal Software Engineer, Duolingo (Confidence: Medium — verbatim, vendor-published customer story)
Confluent's Glean case study reports a narrower but more concrete number: support engineers cut 5–10 minutes off investigation time per ticket (Confidence: Medium — vendor-published, figure confirmed verbatim). None of these are independently audited. All of them describe the same underlying pattern — employees querying an internal RAG layer instead of searching a wiki or a CRM by hand.
Part 2 — The schema markup payoff doesn't survive contact with a controlled test
The WordLift study is the piece of evidence the SEO industry needed and, until now, didn't have. Four domains — editorial, legal, travel, e-commerce — 349 queries tested across seven conditions, 2,443 total evaluations (2,439 valid), retrieval handled by Vertex AI Vector Search 2.0, agentic reasoning via Google's Agent Development Kit, generation by Gemini 2.5 Flash, and judging by Gemini 3.0 Flash as an independent evaluator (Confidence: High — primary source methodology, arXiv:2603.10700v1). Adding JSON-LD/Schema.org markup alone moved accuracy from a 3.62 baseline to 3.89 — a real difference statistically (p_adj=0.024) but a small one by effect size (Cohen's d=0.18) (Confidence: High). The paper's authors are blunt about what that means for markup sitting in a hidden script block on an otherwise unremarkable page:
"[Structured data embedded in hidden script blocks] provides no measurable benefit in flat-text RAG systems." (Confidence: High — verbatim, primary source)
The intervention that actually worked was not adding more markup. It was rewriting the same underlying facts as "enhanced entity pages" — legible, entity-dense prose a retrieval system can parse the way it parses any other page of text. That produced a 29.6% accuracy gain under standard RAG and 29.8% under agentic RAG (Confidence: High) — roughly 160 times the effect size of markup alone, in the same study, on the same underlying content.
Set against that, the claim repeated across the SEO content-marketing ecosystem — "schema markup gets you cited 2.5x more by AI systems" — doesn't hold up when traced. It appears, uncited, on at least seven SEO blogs and content-mill sites; the sites don't even agree with each other, with one citing 2.5x and another citing 3.2x for what is presented as the same statistic (Confidence: High — confirmed via direct multi-site search, no primary study locatable). It is the SEO industry's version of a chain letter: repeated often enough, by enough outwardly authoritative-looking sites, that it reads as settled fact. It isn't.
llms.txt shows the identical pattern at a different layer. An SE Ranking study of roughly 300,000 domains found only 10.13% adoption of the file — 10.54% among mid-traffic sites, just 8.27% among high-traffic sites — and, more importantly, that removing llms.txt as a variable from a citation-frequency prediction model improved the model's accuracy, meaning the file adds noise rather than signal to predicting whether a page gets cited (Confidence: Medium — primary page fetched and figures confirmed verbatim). No major AI system — not Google, not Perplexity, not OpenAI, not Microsoft — has published support for using llms.txt in retrieval or ranking. Two different machine-readable-file interventions, two different studies, the same conclusion: markup and metadata files are not what these systems are actually reading for.
Note for design/production: a two-column diagram works well here — left column "What we told clients to build" (JSON-LD, llms.txt, FAQPage schema), right column "What the controlled study measured" (Δ+0.17/d=0.18 for markup vs. +29.6–29.8% for legible entity pages), with a single arrow showing the effect-size gap. Keep the numbers as the headline; the visual metaphor should not editorialize beyond the data.
Part 3 — Consumer AI search and enterprise RAG have the same failure mode
The instinct is to treat "optimizing for ChatGPT/Perplexity" and "optimizing for internal Copilot/Glean visibility" as separate disciplines with separate playbooks. The WordLift study and the Copilot Studio architecture together suggest they aren't. Copilot Studio's own pipeline retrieves the top three results per knowledge source and then summarizes with citations (Confidence: High) — mechanically, that is retrieval-then-synthesis over text, the same operation WordLift's testbed ran against public web content. A SharePoint page written as a dense, jargon-heavy internal memo is not more legible to that retrieval step than a public case-study page cluttered with hidden schema and thin visible copy. The same content-legibility failure that costs a company citations in a consumer AI Overview also costs an employee a useful answer inside their own company's Copilot deployment.
This matters more, not less, because of where B2B organic traffic is reportedly heading. KEO Marketing's own analysis — undisclosed methodology, described only as "industry analysis across 50,000 B2B websites," and published through Search Engine Land/MarTech, which are the same publisher (Third Door Media), so this is not independently corroborated by a second outlet — claims 73% of B2B websites saw significant organic traffic loss between 2024 and 2025, averaging a 34% year-over-year decline (Confidence: Medium — traced to a real primary source, but treat as a company's own marketing claim, not an audited study). Directionally that lines up with what every enterprise RAG deployment is built to do: intercept the question before it becomes a search, and answer it inside the tool the employee or buyer is already in. Whether or not the 73% figure itself is rigorous, the mechanism it's pointing at — content getting consumed by an intermediary layer instead of clicked through to — is the same mechanism Copilot's and Glean's growth curves describe from the inside.
Part 4 — Legible entity pages are a production discipline, not a markup checkbox
What WordLift's "enhanced entity page" condition actually did was restructure the same facts as clearly labeled, self-contained prose: explicit entity names, explicit relationships between them, explicit context, written so a retrieval system doesn't have to infer structure from a flat block of marketing copy or dig it out of a <script type="application/ld+json"> tag that never gets rendered as visible text at all. That is a content-production requirement, not a technical add-on a developer bolts on after the page ships. It means B2B teams need to write comparison pages, integration docs, and case studies as if a machine will read only the visible prose and nothing else — because, per this study, that is close to true.
Microsoft's own Work Trend Index data offers a structural parallel worth taking seriously here: organizational factors — how a company structures and rolls out AI tools — account for more than twice the measured productivity impact of individual usage factors, 67% versus 32% (Confidence: High — primary Microsoft WorkLab report, based on trillions of anonymized Microsoft 365 signals plus a 20,000-worker, 10-country survey fielded Feb–Apr 2026). The parallel to content strategy is direct: how an organization structures its knowledge — for both external AI search and internal RAG tools ingesting the same case studies and documentation — will matter more than which specific AI tool it adopts or which markup vendor it buys.
What everyone is missing
The industry optimized for the wrong layer of the stack. Schema markup is invisible to the reader and, per the strongest controlled evidence available, nearly invisible to the retrieval system too. Two years of "add JSON-LD for AI visibility" advice were built on a statistic — the 2.5x citation claim — that traces to nothing.
Enterprise RAG collapses the line between "external AI search visibility" and "internal knowledge management." Copilot Studio's retrieval architecture and WordLift's public-web testbed are doing structurally the same operation on text. A company's case studies, docs, and comparison pages now have two audiences that need the identical fix: outside LLMs deciding whether to cite you, and an enterprise assistant like Copilot or Glean deciding whether to surface your content to someone else's employee evaluating your product.
Vendor case studies are the wrong evidence to build strategy on, and they're also the evidence that's actually available. Glean's Duolingo and Confluent stories are self-reported and promotional — genuinely useful directional signal, not proof. The industry needs more studies built like WordLift's: controlled conditions, named models, disclosed effect sizes, replaces-a-cited-figure-with-a-checked-one rigor. That kind of study is rare precisely because it's expensive and unflattering to run.
The llms.txt and schema-markup stories are the same story told twice. Both are low-effort, high-narrative artifacts — a file, a tag — that promise AI visibility without touching the actual writing. Both have now been directly tested and found to add negligible or negative predictive value. The pattern should make anyone pitching the next quick-fix file format or markup standard immediately suspicious of unverified before-and-after case studies.
Future predictions
- Expect more controlled academic studies (arXiv-style, not vendor blog posts) testing specific GEO tactics against RAG accuracy over the next 12 months, driven by exactly the credibility gap this study exposes (Confidence: Medium — reasonable extrapolation from current research momentum, not a confirmed pipeline of specific papers).
- Enterprise RAG vendors (Microsoft, Glean, and others) will likely publish more seat/usage numbers at each earnings cycle as a competitive signal, given the magnitude of the 15M→20M jump already disclosed (Confidence: Medium — based on established quarterly disclosure pattern, not a confirmed future commitment).
- B2B content teams that already invested heavily in schema markup as a primary AI-visibility tactic will face internal pressure to justify that spend once studies like WordLift's circulate more widely (Confidence: Low — directional judgment, not sourced to any disclosed industry reaction).
- "llms.txt adoption" will likely keep rising modestly from its 10.13% baseline for signaling/compliance reasons even without demonstrated citation benefit, mirroring how many low-value technical SEO artifacts persist past the point of proven usefulness (Confidence: Low — pattern-based inference, not a forecast tied to disclosed vendor roadmaps).
- Expect enterprise buyers to increasingly evaluate vendors partly on how well their public content performs inside RAG-based procurement research tools, since Copilot- and Glean-style assistants are already summarizing vendor comparison content for internal stakeholders (Confidence: Medium — consistent with documented Copilot Studio retrieval-and-summarization behaviour, though the specific buyer-behaviour claim itself is not separately sourced).
Practical takeaways
- Stop treating JSON-LD/Schema.org markup as a primary AI-visibility lever. Per the WordLift study, its measured effect size is small (d=0.18); budget and editorial time are better spent on the Entity & Knowledge Architecture of the visible page itself.
- Rewrite core B2B pages — comparisons, integration docs, case studies — as legible, entity-dense prose that names relationships explicitly, rather than relying on hidden markup to carry that structure. This is the intervention that produced the ~30% accuracy gain in the controlled test.
- Don't adopt llms.txt as a visibility strategy. The best available data (10.13% adoption, no positive correlation with citation frequency) says it isn't doing what its proponents claim; treat it as, at most, a low-priority compliance artifact.
- Audit whether your case studies and comparison content are actually retrievable and summarizable the way Copilot Studio's top-3-per-source retrieval step would consume them — this is a Technical Foundations question as much as a copywriting one.
- Build measurement into your GEO program rather than trusting circulating stats: track your own citation rates across AI systems over time using LLM Visibility Monitoring instead of assuming a tactic works because a blog post says so.
- Treat vendor-published case studies (Glean's, Microsoft's, anyone's) as directional signal only — useful for spotting a real underlying mechanism, not as a substitute for a controlled test when deciding where to spend content-production budget.
