A lot of technical SEO advice in 2026 boils down to a single checklist item: "add an llms.txt file", usually alongside a promise that it will help ChatGPT, Perplexity or Google find and quote your site. We found that 90 of the 197 live websites we read for our own AI SEO audit, 45.7%, already had one. Having one made no measurable difference to how much of the site an assistant could quote from. Here's what the file is, who has said they read it, and what our own data adds to the picture.
What is llms.txt?
llms.txt is a plain markdown file at a website's root, yoursite.com/llms.txt, meant to give a language model a short summary of a site plus links to the pages worth reading. Jeremy Howard of Answer.AI proposed the format on 3 September 2024, in a post on answer.ai, reasoning that a model's context window is too small for a whole site and that raw HTML, full of navigation and adverts, wastes what little space there is. The pitch was a robots.txt or sitemap.xml for language models: one small file, deliberately plain, that points a model at the pages that matter.
Does any AI assistant read it?
None has confirmed that it does. Google's own Search Central guidance, published 15 May 2026 as Optimizing your website for generative AI features on Google Search, states it directly: "you don't need to create new machine readable files, AI text files, or markup to appear in these features." No equivalent commitment exists from OpenAI, Anthropic or Perplexity either. llms.txt remains a proposal by one researcher, not a standard adopted by any of the assistants marketers write it for.
Ahrefs checked what happens to the file in practice. According to its study of 137,210 domains with May 2026 traffic, 28% published an llms.txt, and 97% of those files received zero requests that month. Of the roughly 1,100 domains whose file was fetched at all, AI search bots, OAI-SearchBot, PerplexityBot and Claude Web among them, made up just 1.1% of the requests. Most of the traffic the file did get came from SEO audit tools, unidentified bots and crawlers like Googlebot checking what was there, not from an assistant retrieving an answer.
What our audit of 197 sites found
Our own numbers say the same thing from a different angle: having the file doesn't change what a site's pages offer an assistant. We scored every section of every site we read on four things, evidence being one: does the section carry a number, a named source or a link a model could quote. Adopters and non-adopters landed within noise of each other.
| Group | Sites | Evidence score (median) |
|---|---|---|
Publishes llms.txt | 90 (45.7%) | 25.0 |
No llms.txt | 107 (54.3%) | 19.5 |
| Our own research, 197 sites read 31 July-6 September 2026. |
The gap is not statistically significant (Mann-Whitney, p = 0.18, one reading per site). It's also not a one-off: an earlier read of 128 sites in August found the same shape, 25.7 against 23.3, p = 0.45. Two separate corpora, two null results.
The file itself explains why. The median site that publishes an llms.txt still has half its sections under 40 words and three quarters carrying nothing quotable, the same thin, unevidenced writing we find on sites without one. The file is a pointer to pages, not a rewrite of them, and pointing at a weak page doesn't strengthen it.
Why the file doesn't change what gets cited
An assistant that cites a source is quoting a specific passage, not crediting a whole domain. Wherever it found that passage, on a normal page or through a link in an llms.txt summary, it still needs a short, self-contained section with something checkable in it: a figure, a quotation, a named source. Adding a pointer file doesn't add evidence to the page at the other end of the link. That's the gap our numbers are measuring: not whether the file gets read, but whether reading it would have found anything worth quoting if it did.
What to check instead
Our own free AI SEO audit reports whether a site has an llms.txt, because it's a fair question to ask, but doesn't score it for exactly this reason: it's "a convention no major crawler is confirmed to read." What the audit does score is the part our data shows correlates with getting quoted:
- Access: can a named crawler, Googlebot, OAI-SearchBot, PerplexityBot and the rest, fetch the page at all, past
robots.txtand any firewall rule. - Rendering: is there readable text before JavaScript runs. In the same 197-site read, 21.3% of sites had at least one page that stayed blank, or partly blank, until JavaScript executed.
- Discovery: does the sitemap lead to the pages worth reading.
- Chunking: does each section stand alone, so it still makes sense when an assistant lifts it out of the page around it.
- Evidence: does the section carry a number, a quotation or a link, the one dial our llms.txt comparison shows genuinely divides sites.

The report above shows the shape of what comes back: one score out of 100, then a breakdown by dial, worst finding first with the exact page it's on. Illustrative, this one shows 84/100 for a sample company with 842 pages read, access scoring a clean 100 while length and evidence trail at 49 and 36, alongside three named fixes such as pages sitting empty until JavaScript runs. llms.txt doesn't appear in the score at all, on this site or any other, because none of the four dials depends on it.
If you already have one, keep it short and accurate
Nothing in our data says an llms.txt file costs a site anything. The problem is what it tends to replace on a team's to-do list, not what it does once it's live. Write it as a short, factual summary: what the company does, in plain sentences, plus links to the pages a model would actually need, pricing, product pages, documentation. Don't pad it with marketing language a crawler can't verify, and don't treat writing it as a substitute for fixing the pages it links to. Our own llms.txt is a reasonable shape to check yours against. A file that points to a section with no number, no named source and nothing to quote hands a model a route to a page that still has nothing on it worth citing.
What this means if you're deciding whether to add one
Keep an existing llms.txt if you already have one; there's no evidence it costs you anything to leave in place. But don't treat writing one as the fix for a brand that isn't showing up in AI answers, and don't let it be the only line item on the technical AEO checklist. Time spent shortening sections to stand alone, adding a citable figure or source to the ones that currently have neither, and confirming a crawler can reach and render the page moves the dials our data shows separate sites that get quoted from sites that don't. A pointer file to pages with nothing quotable on them is still a pointer to pages with nothing quotable on them. The metrics behind AI search visibility are the place to start, and how AI crawlers find and cite a brand covers the access and discovery side in more depth.


