What is llms.txt, and do I actually need one?
A 2024 proposal for a Markdown index file, what Google has said about it, and why crawl access matters more than the file itself.

Short answer: llms.txt is a proposed Markdown file, published in September 2024 by Jeremy Howard and hosted as an open specification, meant to give AI agents a short, curated index of a site's most important pages. It is not a Google standard, not required for AI Overviews, and not confirmed by OpenAI, Google or Anthropic as something their assistants read before answering.
The idea is reasonable on its own terms. A model working through a limited context window benefits from a short table of contents pointing at the pages worth reading, instead of wading through navigation and cookie banners on every page. Where the proposal gets oversold is the jump from a sensible idea to a ranking requirement, and a lot of the content selling llms.txt implementations makes exactly that jump.
What follows is what the file actually contains, what the platforms have said about it, and where the same effort is better spent first.
At a glance
An independent proposal, not a standard
Jeremy Howard published the llms.txt specification in September 2024 on GitHub, where it still lives. No standards body, not W3C, not IETF, has adopted it, and no browser or search engine requires it.
Google says you don't need it
Google's Search Central guidance on AI features states plainly that no new machine-readable files, AI text files or markup are required to appear in AI Overviews or AI Mode, and no special schema.org markup either. llms.txt falls squarely in that category.
No major engine confirms reading it
OpenAI's published crawler documentation covers OAI-SearchBot, GPTBot and ChatGPT-User in detail and says nothing about parsing an llms.txt file. Anthropic and Google have not confirmed it either. Some agent frameworks and RAG tools do read it; the large consumer assistants have not said they do.
Crawl access decides whether any of this matters
A carefully written llms.txt on a site that blocks GPTBot and OAI-SearchBot in robots.txt changes nothing. Check what robots.txt actually allows before spending time curating a page list for crawlers that can't reach the pages anyway.
What the file is supposed to do
llms.txt sits at the site root, written in Markdown, and lists the pages an AI agent should read first, usually with a one-line description of each. A companion convention, llms-full.txt, concatenates the full content of those pages into one file so an agent can pull everything in a single request instead of following links one at a time.
The pitch is that an agent working inside a limited context window gets more value from a hand-picked index than from crawling an entire site and hoping the important pages surface. That's a real problem for very large documentation sites, which is why Mintlify, GitBook and several other docs platforms generate the file automatically.
It's a different job from robots.txt, which controls permission, not curation, and from sitemap.xml, which lists every URL for discovery, not the handful worth reading closely. Confusing the three is the most common mistake in advice about this topic.
What Google and the AI engines have actually said
Google's own AI-features documentation is the most direct statement on record: no new machine-readable files, AI text files or markup are needed to appear in AI Overviews or AI Mode. The advice underneath is the same advice Google has given for two decades, be crawlable, be helpful, be fast.
OpenAI documents four crawlers on its developer site: OAI-SearchBot, GPTBot, ChatGPT-User and OAI-AdsBot, and describes exactly what each one is for. None of that documentation mentions llms.txt. Anthropic and Perplexity have not published anything confirming their assistants fetch it at answer time either.
That silence doesn't make the file harmful. It means the honest claim is unconfirmed, not proven, and any page promising it will get you cited faster is asserting something none of the companies involved have actually said.
Where the same effort is better spent
Before curating a page list for crawlers, confirm the crawlers can get in at all. Open robots.txt and check it allows GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and Google-Extended, unless blocking one of them was a deliberate choice. A single Disallow: / line left over from a staging build blocks everything downstream regardless of what llms.txt says.
Then check what those crawlers actually receive. Several of them don't execute JavaScript. If pricing, product details or service descriptions only render after scripts run, an AI crawler sees a mostly empty page no matter how complete the content looks in a browser.
Only once both of those are solid does a curated index add anything. Treat it as a small, cheap addition once the fundamentals hold, not a first move.
Does LumiRank publish one?
Yes. This site serves a Markdown index at /llms.txt and a full mirror at /llms-full.txt, both generated from the same service, case-study and article data that renders the HTML pages, so the two can't drift out of sync with each other. We built it that way because a hand-maintained second copy of a site is the version that goes stale first.
We don't expect it to move rankings or AI citations on its own, no engine has confirmed reading it. It sits next to working robots.txt access and JavaScript-independent content, which is where the real visibility work happens.
Three files, three different jobs. Mixing them up is the most common misunderstanding in coverage of this topic.
| File | What it controls | Who reads it |
|---|---|---|
| robots.txt | Which crawlers may fetch which URLs | Every compliant crawler: Googlebot, Bingbot, GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot and more |
| sitemap.xml | Which URLs exist and when they last changed | Search engine indexers, to find and prioritise pages to crawl |
| llms.txt | A curated Markdown index of pages worth reading first | Not officially confirmed by any major consumer AI assistant; some agent frameworks and RAG tools do read it |
robots.txt and sitemap.xml are decades-old, universally honoured standards. llms.txt is a 2024 proposal with no such guarantee. Get the first two right before spending time on the third.
What is llms.txt, and do I actually need one?: common questions
Is llms.txt the same thing as robots.txt?
No. robots.txt controls which crawlers may access which URLs, and every major crawler honours it. llms.txt is a curated Markdown index of recommended pages, and no major AI engine has confirmed it reads that index before answering. Get robots.txt right first; it's the one of the two with a guaranteed effect.
Will adding an llms.txt file get my pages cited by ChatGPT or Perplexity?
There's no evidence either way from the companies themselves. OpenAI's crawler documentation doesn't mention it, and neither does Google's or Anthropic's public guidance. Some non-consumer agent frameworks and RAG pipelines do read it. Treat any specific claim about ChatGPT or Perplexity reading it as unverified, and check their current documentation before acting on it.
What should an llms.txt file contain if I add one?
A short Markdown page listing your most important URLs with a one-line description of each, grouped by section, following the format published at llmstxt.org. Keep it accurate: list pages that genuinely explain what the business does, and update it when those pages change. An index pointing at outdated or thin pages is worse than no index.
Is there a downside to publishing one?
Not much. It takes little effort to build and doesn't compete with anything else on the site for attention. The real risk is treating it as a substitute for the fundamentals: crawl access, content that renders without JavaScript, and pages that actually answer the question. An llms.txt file can't fix any of those three.
Where ChatGPT gets its answers
Does schema markup help SEO?
Getting into AI Overviews
Related services
+3
Further reading from the journal
+2
Want to be the brand the models name?