Follow up to the robots.txt pull I did earlier this week. Same 109 domains, this time checking whether they actually publish an llms.txt, and while I was in there, ai.txt as well.
On method, because it changes the number: I checked content type and body rather than counting any 200 response. Plenty of sites return a styled 404 with a 200 status, and if you count those you'll massively overstate adoption.
43 of the 109 have an llms.txt. Sounds like real traction until you break it out.
SaaS 83% Ecommerce 40% Finance 40% Marketing and SEO publishers 26% Travel 22% Health 12% News 4%
Twenty five of thirty SaaS companies have one. One of twenty two news sites does. That isn't an adoption curve, that's two unrelated populations.
The SaaS number stops being impressive once you open the files. Stripe's llms.txt is a markdown index of their documentation. So is Zendesk's. These companies already had docs in markdown with clean URLs, so shipping an LLM readable index cost them roughly an afternoon. It's a docs artefact that happens to live at the root, not a marketing decision anybody agonised over.
Two things I'd want to know before paying somebody to write one.
Four of the six ecommerce files aren't decisions at all. Allbirds, Glossier, Casper and Gymshark all serve the same generated file, opening with "# Agent Instructions" and the brand name, describing how agents can interact with the store. Shopify produces it. I doubt anyone at those companies wrote it or knows it's there. So some portion of any adoption stat you read is platform defaults rather than intent.
And nobody agrees what the file is for. Stripe and Target use theirs as a machine readable content index, effectively a sitemap for models. TIME uses theirs to say the opposite. Their file is a policy notice with a contact address and "disallow: *" sitting in it. Same filename, completely opposed intent, and nothing enforces either reading.
ai.txt came back zero out of 109. Not one site. That format is dead and you can stop reading the posts that still recommend it.
Content-Signal directives in robots.txt showed up on 10 of 109, and six of those ten are marketing or SEO publishers. That's the same group that had the sharpest crawler policy in the previous pull. They're consistently ahead of every other vertical on this, which figures, since they're the people writing about it.
The honest limit of all this: I can't tell you whether any of it affects citations. I have no way to measure whether a site with llms.txt gets cited more often than one without. Someone asked me exactly that on the last post and the real answer is I don't know, and I'd want to see the controlled test before believing anyone who says they do.
So where I land. If your docs are already in markdown, publishing one takes an afternoon and costs nothing, go ahead. If somebody's charging you real money for it, ask what evidence they have that anything reads it. At the moment it's a convention with adoption concentrated in a single vertical, partly generated by platforms rather than chosen, with two incompatible uses in the wild and no demonstrated effect on anything measurable.
Source: r/DigitalMarketing · by /u/Conscious-Market8982