Is llms.txt Worth Shipping in 2026? What 137,000 Sites Showed
Ahrefs found 97% of llms.txt files get zero AI-crawler traffic and no major engine has confirmed reading one. The file still has exactly one job worth doing.

Companion film
97% of llms.txt files, zero crawler traffic
The evidence against llms.txt in one pass, and the one use case that survives it.
1:15captionsthe same scan as the screenshots below
Read the transcript insteadHide the transcript
0:01Should you publish an llms.txt file? Ship it if it takes 20 minutes. Never count it as a visibility win.
0:09Ahrefs looked at about 137,000 sites. Ninety-seven percent of those files got zero AI crawler traffic. Not low. None.
0:19No major provider has committed to reading it. Google says it doesn't support the file and isn't planning to, and one of its engineers compared the idea to the old keywords meta tag.
0:31It does have one real use. Not as a broadcast signal, but as a knowledge file you hand to an assistant yourself. Paste it into a project, attach it to a custom assistant. There, it works.
0:43The reason to be blunt about it is that it feels productive. And that feeling is expensive when it replaces work with a measured effect.
0:52Check which AI crawlers you actually allow. Earn third-party coverage, because most AI brand mentions come from someone else's page. And add statistics and quotations: 41% and 28% more visibility in the Princeton study.
1:07Ship the cheap hedges. Label them as hedges. Fourteen free tools, at whereamimentioned.com.
The short answer: publish it if it takes you twenty minutes, never count it as a visibility win, and never let it displace something that works. The long answer is why that's the only defensible position given what's been measured.
What the 2026 data actually shows
Ahrefs analysed roughly 137,000 sites and found 97% of llms.txt files received zero AI-crawler traffic in May 2026. Not low traffic: none.
A separate look at over 515 million LLM bot traffic events found the share of requests touching /llms.txt statistically negligible against total crawl volume.
No major provider has committed to it. As of early 2026, OpenAI, Google, Anthropic, Meta and Mistral had all declined to publicly confirm reading or acting on llms.txt in production systems. Google has been the most explicit: it does not support the file and is not planning to, with John Mueller comparing the idea to the long-discredited keywords meta tag on the grounds that both are what a site owner claims about itself, checkable only against the page you would have to read anyway.
Independent tests found no citation effect. SE Ranking modelled 300,000 domains and found the file carried no signal: dropping it as a feature made the model more accurate, not less. Trakkr compared 37,894 domains and measured 6.8 citations for sites with the file against 6.7 for sites without, p = 0.85.
That's an unusually clean evidence picture for this field. Most GEO tactics suffer from an absence of data. This one has data, and the data says the mechanism people hope for isn't operating.
Why the idea was reasonable
llms.txt was proposed to solve a real problem. Language models have finite context windows, HTML is mostly navigation and markup with content somewhere inside it, and JavaScript-rendered pages are expensive or impossible to parse. A single markdown file at a predictable path (a title, a one-paragraph summary, curated links grouped by section) is a sensible thing to want.
The proposal was also modelled on the two conventions that did succeed, robots.txt and sitemap.xml, which is why it feels intuitive. But both of those succeeded because the consumers adopted them first and publishers followed. llms.txt has the sequence backwards: publishers adopted it enthusiastically and the consumers never arrived. A standard nobody reads is a file.
The one use case that is real today
llms.txt does work, just not as a broadcast signal. It works as a knowledge file you hand to an assistant yourself.
Concretely: paste llms-full.txt into a Claude Project, attach it as a custom GPT knowledge file, or drop it into a chat when you want an assistant to answer questions about your product accurately. In that flow the file does exactly what it was designed to do, because you're the consumer and you definitely read it.
The adjacent proven use is developer tooling. AI coding assistants including Cursor and GitHub Copilot fetch documentation files like these to pull clean docs into context. If you publish an API or a developer product, that's a genuine reason to maintain one, and it has nothing to do with brand visibility.
Notice what both real use cases have in common: a specific consumer you can name. That's the test for any GEO tactic. "Engines might pick it up" isn't a mechanism.
Ship one in twenty minutes, then stop thinking about it
If the cost is genuinely twenty minutes, the expected-value calculation is fine. Three free tools do the work:
- The llms.txt generator crawls your site, discovers pages via sitemap or link traversal, and produces the index format: H1 title, blockquote summary, section-grouped links.
- The llms-full.txt generator inlines the actual page content into one markdown file. This is the version worth having, because it's the one you can paste into an assistant.
- The llms.txt validator checks an existing file's structure and flags the problems that make a file useless even to a willing reader.


Every one of those pages carries the same caveat this article does, because a tool that oversells its own output is worse than no tool.
Where that twenty minutes would go further
The reason to be blunt about llms.txt has nothing to do with harm. It's that the file is satisfying: a discrete, shippable artifact that produces the feeling of having done AI optimisation. That feeling gets expensive when it substitutes for work with a measured effect.
Ranked by evidence, here's what the same afternoon buys instead:
1. Check which AI crawlers you actually allow. Blocking a retrieval crawler removes you from the retrieval index: a direct, mechanical loss of citations, not a speculative one. Published estimates of how many sites are doing this range from about 5% to about 45% depending entirely on which population got sampled, so ignore the headline figures and read your own file, which takes a minute. The user agents are not interchangeable: blocking GPTBot has no effect on OAI-SearchBot, and blocking ClaudeBot has no effect on Claude-SearchBot. The crawler access checker reads your robots.txt against the current list; the article on it explains which agent does what.
2. Earn third-party coverage. AirOps traced roughly 85% of brand mentions in AI answers to third-party pages rather than the brand's own domain. In our own demo scan, one Iceland Review feature was cited five times while the brand's own website was cited twice. One earned article outperformed the whole site.
3. Add the content elements with measured lift. The Princeton GEO study found that adding statistics raised position-adjusted visibility by around 41%, adding quotations moved the subjective-impression metric by around 28%, and citing sources lifted a page sitting fifth by over 100%, while keyword stuffing was among the worst performers. Lower-ranked pages gained the most. Two caveats the write-ups usually drop: the experiment ran on GPT-3.5-turbo against a fixed set of five competing sources, so every figure is a gain relative to four rivals in a closed contest rather than a forecast for a live grounded answer, and the optimisations were written by a model, not by a person. Even discounted for both, it is the only tactic in this list with a published effect size attached.
4. Make your answers extractable. A crisp definition, a comparison table and a checklist out-cite a long essay. The hard number people quote here is BrightEdge's 82.5% of citations coming from deep pages rather than homepages, and it is worth knowing that it measures Google AI Overviews, a surface this product does not probe and will not claim to. Treat it as evidence about Google. The edit it argues for is a structural one to pages you already have, which is cheap either way.
What to say when someone asks whether you have llms.txt
Say yes, you publish one, it took twenty minutes, and no, it isn't why your citations moved. Then show them the crawler access report, the earned coverage, and the citation share per engine.
Why this matters beyond one file: llms.txt is the current test case for whether a team can hold a tactic and its evidence separately in mind. Schema markup sits in the same category, genuinely useful for entity disambiguation and for the index that feeds AI Overviews, and no kind of direct citation lever whatever the tooling implies. So does the FAQ markup on this very page. Google deprecated FAQ rich results for every site in May 2026, so this markup earns no expandable results anywhere. It ships because the question-and-answer format is what assistants quote, and because valid machine-readable Q&A costs nothing.
Ship the cheap hedges. Label them as hedges. Spend the real budget where something has been measured.
FAQ
Does llms.txt improve AI citations? There's no evidence that it does. Independent analyses from SE Ranking and Trakkr found no link between having the file and being cited more often, and Ahrefs found 97% of llms.txt files receive no AI-crawler traffic at all.
Do any AI engines read llms.txt? None of the major providers has publicly committed to reading it in production. Google has said it doesn't support the file and has no plans to. Some smaller engines and developer tools do fetch it, which is a real but narrow use case.
Should I publish llms.txt anyway? If it costs you twenty minutes and you never report it as a visibility win, yes: it's a cheap hedge with a genuine present-day use as a knowledge file for assistants. If it displaces work on crawler access, content structure or third-party coverage, no.
What is the difference between llms.txt and llms-full.txt? llms.txt is an index: a title, a summary, and curated links grouped into sections. llms-full.txt inlines the actual page content into one large markdown file. The full version is the more useful of the two today, because pasting it into an assistant hands that assistant your whole site as context.
Corrections and updates
What changed after publication, and why. An article that argues for showing your working does not get to edit itself quietly.
- Corrected the Princeton GEO figures. The 41% belongs to adding statistics, not quotations; quotations moved a different metric by about 28%; citing sources lifted a fifth-placed page by over 100%. Added the conditions the original omitted (GPT-3.5-turbo, five competing sources, model-written edits). Attributed the 82.5% deep-page figure to BrightEdge and scoped it to Google AI Overviews. Removed an unsourced claim that 41% of B2B sites block a major AI crawler. Linked every remaining external number. The companion film repeated both mistakes on screen and in its voiceover, so it was re-recorded and re-rendered rather than left to contradict the article.