- Ahrefs checked 137,210 domains and found 97 percent of the llms.txt files on them received zero requests during May 2026.
- SE Ranking tested roughly 300,000 domains for a link between llms.txt and AI citations. Removing the file from their model made the model more accurate, not less.
- Across both studies that is more than 437,000 domains examined, with no measured benefit in either.
- Of the requests that do reach an llms.txt, 77 percent come from bots that are not AI tools at all. SEO audit crawlers alone account for 21.7 percent.
- Google's documentation states directly that no machine readable file, AI text file or special markup is needed to appear in AI Overviews or AI Mode.
- What did correlate with AI citations in Ahrefs' separate work was branded web mentions at a Spearman coefficient of 0.664, plus content freshness.
- Publishing the file is cheap. Believing it does something is what costs you, because it displaces the work that measures.
Ninety-seven percent. That is the share of llms.txt files that went completely unread during May 2026, measured across 137,210 domains with real traffic. Not lightly read. Not read by the wrong bots. Zero requests, for a file whose entire purpose is to be requested.
The number comes from Ahrefs' study of 137K sites, which checked each domain root for a file returning HTTP 200, then pulled the actual request logs for the /llms.txt path and sorted them by user agent. It is the kind of study that is boring to run and hard to argue with, because it measures fetches rather than opinions.
If you have been carrying llms.txt on a to-do list, this piece is the permission to close it. If you already shipped one, nothing bad happened, and you can stop wondering whether it is working. What follows is what the data actually says, who is really reading the file, and where the same hour of work produces something measurable.
What was llms.txt supposed to do?
The proposal is reasonable on its face. Put a markdown file at your domain root listing your most important pages, so a language model fetching your site gets a curated map instead of crawling a navigation menu, a cookie banner and a footer. Think robots.txt for meaning rather than permission, or a sitemap written for a reader rather than a parser.
The idea spread fast because it costs almost nothing to adopt and it feels like insurance. Nobody wanted to be the site that was invisible to AI answers because they skipped a text file. Adoption reflects that: Ahrefs found 28 percent of the domains it studied publish one, which is more than one site in four, for a convention no major AI platform ever committed to reading.
That gap between adoption and commitment is the whole story. The file was widely implemented on the assumption that someone would eventually consume it. Two studies have now checked whether anyone does.
What did the data actually show?
Two independent teams asked slightly different questions and got the same answer from opposite directions. One measured whether the file is fetched. The other measured whether having it changes anything. Neither found what proponents hoped for.
| Study | Sample | What it measured | Headline finding |
|---|---|---|---|
| Ahrefs, published 2026 | 137,210 domains with traffic in May 2026 | Actual HTTP requests to /llms.txt, classified by user agent | 97 percent of valid files received zero requests in the month. Roughly 1,100 domains absorbed all measured traffic. |
| SE Ranking, published 7 November 2025 | Nearly 300,000 domains | Statistical correlation plus an XGBoost model predicting AI citation frequency | Adoption at 10.13 percent, and dropping llms.txt from the model improved its accuracy. |
| Combined | More than 437,000 domains | Both fetch behaviour and citation outcome | No measured benefit on either axis. |
| Google documentation | Platform guidance, not a study | What a site must publish to appear in AI features | No machine readable file, AI text file or special markup is required. |
The SE Ranking result deserves a second look because it is stronger than a null finding. As Search Engine Journal reported on the study, the team ran an XGBoost model over domain level citation frequency and found that removing the llms.txt variable improved the model's accuracy. A feature that makes a predictive model worse is not neutral. It is noise, and the model was better off without it.
The original SE Ranking write-up adds a detail that undercuts the survivorship argument people reach for: adoption was flat across traffic tiers, running 9.88 percent for the smallest sites, 10.54 percent in the middle, and 8.27 percent among sites with more than 100,000 visits. The biggest sites are slightly less likely to bother, not more.
Who is actually reading llms.txt?
Mostly tools that audit websites for a living. This is the part of the Ahrefs study that gets skipped, and it is the most useful part, because it explains why anyone ever thought the file was working.
Of the small share of files that did receive requests, 96 percent of those requests came from bots, and 77 percent of the bots involved were not AI tools at all. Here is the breakdown Ahrefs published, reordered by size.
| Requester category | Share of requests | Is it an AI system consuming your content? |
|---|---|---|
| SEO audit tools | 21.7 percent | No. Your own auditing stack, or a competitor's. |
| Other or unidentified | 14.9 percent | Unknown. |
| General web crawlers | 13.1 percent | No. Conventional indexing. |
| Tech profiling tools | 11.6 percent | No. Stack detection. |
| AI agents and agentic infrastructure | 10.5 percent | Partly. Coding agents fetching a file on request. |
| GEO and AEO tools | 5.8 percent | No. Visibility trackers checking the file exists. |
| AI training crawlers | 5.3 percent | Yes, for training rather than answering. |
| llms.txt discoverability bots | 3.6 percent | No. Bots that exist to count the files. |
| AI assistants | 2.5 percent | Yes. |
| AI retrieval bots | 1.1 percent | Yes, and this is the category that would feed an answer. |
Read the last row again. The bots that fetch pages in order to answer a live question, the ones whose attention the whole exercise was meant to attract, account for 1.1 percent of traffic to a file that 97 percent of the time receives none.
There is a second finding in the study that quietly settles the argument: no AI bot ever requested an llms.txt on a domain that did not have one. They do not go looking. A file only gets fetched when something already decided to fetch it, which means the file cannot create discovery, only satisfy it.
The one AI category with real volume is agentic infrastructure at 10.5 percent, driven by coding agents. That is a genuine use case and it is worth naming honestly: an agent working in your repository may fetch your documentation index because a developer asked it to. That is a developer experience win, not a search visibility win. If your audience is developers pointing agents at your docs, the file has a job. Those same agents raise their own questions, which we covered in our look at the injection risks that come with AI coding agents.
Does Google want you to publish one?
No, and it says so in plain language rather than leaving it to inference. Google's documentation on AI features and your website addresses the question head on: you do not need to create new machine readable files, AI text files or markup to appear in AI Overviews or AI Mode, and there is no special schema.org structured data to add. The markup that does earn its keep is the ordinary commercial kind, which is also what the AI shopping agents now arriving on storefronts read before they build a cart.
SE Ranking makes the same point about the other major platform: OpenAI's own guidance directs site owners to manage crawler access through robots.txt, not through a separate content index. Two of the largest consumers of web content have documented what they read, and neither list includes this file.
So what does correlate with being cited?
Things that are harder to do, which is why the file was attractive. Ahrefs ran separate work on what actually moves AI search visibility, and the measured associations look nothing like a configuration file.
| Lever | Measured strength | What it means in practice |
|---|---|---|
| Branded web mentions | Spearman 0.664 across roughly 75,000 brands | How often your name appears on other people's pages, linked or not. |
| Mentions on heavily linked pages | Roughly 0.70 | Where the mention sits matters as much as how many there are. |
| Branded anchor text | 0.527 | Links whose anchor carries your brand name. |
| Mentions on high traffic pages | Roughly 0.55 | Reach of the page carrying the mention. |
| Content freshness | 13.1 percent preference for recently updated pages | Cited content ran 25.7 percent fresher than what ranks organically, across 17 million citations on 7 platforms. |
| Publishing llms.txt | No measurable effect in either study | Nothing. |
Correlation is not causation and these figures do not prove that earning a mention causes a citation. What they do establish is a ranking of where evidence exists at all. One column has coefficients attached. The other has a null result from 437,000 domains.
The freshness number is the most actionable line in the table, because it points at pages you already own. Updating a genuinely useful page that has drifted out of date is cheaper than earning a new mention, and it is measurable in your own logs within weeks.
Why do the two studies disagree on adoption?
They report 28 percent and 10.13 percent for what sounds like the same measurement, and the gap is worth resolving rather than averaging.
The samples are not the same population. Ahrefs looked at domains inside its own Web Analytics product that received traffic during May 2026, which is a set of sites whose owners installed an analytics script and are actively managing them. SE Ranking crawled a broader pool of domains without that filter. A site that someone actively tends is far more likely to have picked up a new SEO convention than a domain sitting in a general crawl, so the higher number describes engaged site owners and the lower number describes the web.
Both readings support the same conclusion from different angles. Among people paying attention, better than one in four adopted the file. Across the web at large, one in ten did. In neither population did adoption produce a measurable outcome, which means the effect is absent at both levels of engagement rather than diluted by inattentive sites.
The discrepancy is also a reminder to check what a percentage is a percentage of before repeating it. Two credible teams, two defensible samples, numbers that differ by a factor of nearly three, and neither is wrong.
How would you know if it ever started working?
By watching your own logs, which is a five minute setup and the only evidence that will ever be about your site rather than someone's aggregate.
Filter your access logs for requests to /llms.txt and group them by user agent over a rolling thirty days. You are looking for three specific names, because they behave differently. Training crawlers fetch broadly and tell you nothing about answers. Assistant and retrieval bots fetch when a live question is being answered, and those are the requests that would matter. Coding agents fetch on a developer's instruction, which is a docs signal rather than a search signal.
Set the bar before you look: if retrieval bots are not fetching the file at all, nothing downstream can be attributed to it. That test costs nothing and it inoculates you against the most common mistake in this area, which is seeing traffic in a bot report, feeling validated, and never checking that the traffic came from an audit tool you pay for yourself.
Ahrefs' finding that no AI bot ever requested the file on a domain that lacked one gives you the control condition for free. These bots do not probe. If yours is being read, something specifically chose to read it, and your logs will name what.
One more number from the Ahrefs freshness work is worth carrying into that log review. The analysis covered 17 million citations across 7 platforms and found that the content assistants cite runs 25.7 percent fresher than what ranks in organic results. Different systems, different appetites. A page that has been stable for three years can hold a search ranking and still be passed over by an assistant answering the same question, which is a reason to date your updates visibly rather than a reason to churn your library. Freshness is only one of the inputs, and the measured content changes that move a citation rate inside AI answers turn out to be more specific than recency.
Is there any reason to publish it anyway?
Three, and none of them are search visibility.
The first is developer experience. If people point coding agents at your documentation, a curated index of your most important pages saves those agents a crawl. The 10.5 percent agentic share in the Ahrefs breakdown is real traffic doing real work. Treat it as a docs feature, put it in your docs backlog, and stop expecting it to show up in a citations dashboard.
The second is optionality. The file costs an hour to generate if your CMS can emit it, and platforms change their minds. SE Ranking's own conclusion allows for this: low effort preparation for a possible future, provided you are honest that the present value is zero.
The third is that a genuinely machine readable interface to your systems is a real category, and llms.txt is a weak member of it. If your goal is for an assistant to work with your product rather than read about it, the protocol built for that job is the one to learn. Our explainer on how the Model Context Protocol connects an assistant to your systems covers the difference, and the practical walkthrough of standing up an MCP server covers what it takes to ship one. That is a file being read because something needs it, which is the property llms.txt never acquired.
If you keep an llms.txt, treat it like code. Ahrefs points out it should be version controlled and reviewed, because a hand-maintained index of your important URLs rots quietly and a stale one is worse than none. Anything auto-generated by your CMS is safer than anything a person edits twice a year.
What should you do this week?
Delete the task, or demote it. Then spend the hour on one of the levers with a coefficient next to it.
Pull your ten highest value pages and check the last meaningful update date on each. Not the timestamp your CMS bumped, the date the content actually changed. Anything over a year old that still gets traffic is a candidate, and refreshing it is the cheapest move on the table given the 13.1 percent freshness preference.
Then look at mentions rather than links. Search your brand name and count how many results are pages you do not control. That number, not a text file, is what the strongest measured correlation is made of. Growing it is slow work involving other people's sites, which is exactly why a configuration file was such an appealing alternative.
The honest summary is that llms.txt asked a good question and got a null answer. The good question was whether sites should publish something machine readable for AI systems. The answer, so far, is that the systems already read your pages, and they told us in their documentation what they need, which is content worth citing on a site that other people talk about. Which moves the work back onto the pages themselves, where the useful question is how much editing generated copy needs before it is worth citing rather than which file you publish alongside it.