---
title: "How many websites have an llms.txt file? 2026 adoption data"
description: "We checked /llms.txt on 5,902 top websites: 11.5% publish a valid file, 2.1% an llms-full.txt, and 15.7% return an HTML page at the address instead."
canonical: "https://underneath.agency/research/llms-txt-adoption-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI crawlers

# How many websites have an llms.txt file? 2026 adoption data

llms.txt is a proposed plain-Markdown file at the root of a website that gives language models a short map of the site. We requested it from every live website in the Tranco top 10,000 on 26 September 2026 and checked whether what came back was a real llms.txt or a web page served in its place. 11.5% of top sites publish a valid file. Even more sites answer the address with an HTML page, which a naive count would score as adoption. A second fetch two days later, with a control address no site publishes, confirmed both figures.

The study measures one thing: whether a site publishes the file. It does not measure whether any AI system requests the file, uses it to find pages, cites the site or uses its content in answers. The section on stages below says what is known about each of those, and from whom.

## The short version

1. 680 of 5,902 live top websites (11.5%) serve a valid llms.txt: HTTP 200, plain text, opening with a Markdown heading as the proposal specifies.
2. 929 sites (15.7%) return an HTML page at /llms.txt, usually a soft 404 or the homepage. For every valid file there are 1.37 of these, so counting any HTTP 200 response more than doubles the apparent adoption.
3. Adoption falls with popularity rank but not by much: 15.2% of the top 1,000 sites, 12.4% of sites ranked 1,001 to 5,000, 10.2% of sites ranked 5,001 to 10,000.
4. Only 2.1% of live sites serve a valid llms-full.txt, the companion file with the full text; 16.3% of llms.txt publishers also publish one.
5. The typical valid llms.txt is 7,706 bytes with 36 links; 80.6% include the one-paragraph summary the proposal recommends, but only 11.9% meet every part of the proposal including links to Markdown pages.
6. The HTML responses are catch-all pages: 96.6% of those sites also answer a made-up address with HTTP 200. None of the 680 valid files is a catch-all response.
7. The measurement is stable: fetched again on 28 September, 99.2% of sites were in the same state and 677 of the 680 valid files were still valid.

## Which stage this study measures

Between a page on a website and a sentence in an AI answer there are several separate steps. A result at one step says nothing about the next, so each claim about llms.txt should name its step.

| Stage | Question | What is known |
|---|---|---|
| Publication | Does the site serve a valid file? | Measured here: 11.5% of live top sites |
| Discovery | Does an AI system request the file? | Not measured here. Ahrefs found most files receive no requests at all |
| Retrieval | Does an AI system use the file to find or read pages? | Not measured by anyone we found |
| Citation | Is the site cited more often because of the file? | Not measured here. SE Ranking found no correlation, which cannot show cause either way |
| Answer use | Does the file’s content change what the answer says? | No evidence either way |

Every finding below is about the first stage.

## What llms.txt is

The llms.txt proposal (llmstxt.org) asks sites to publish a Markdown file at /llms.txt with an H1 title, a short blockquote summary, and sections of links to the pages that matter most, ideally to Markdown versions of those pages. A second file, /llms-full.txt, can hold the full text of the site in one document. The idea is that an AI system can read one small, clean file instead of crawling and parsing a whole site.

## What we measured

For each of the 5,902 domains in the Tranco top 10,000 whose homepage responded with HTTP status below 400, we requested /llms.txt and /llms-full.txt once and classified the response:

| Response at /llms.txt | Sites | Share of live sites |
|---|---|---|
| Valid llms.txt (HTTP 200, text, Markdown H1) | 680 | 11.5% |
| Plain text but no Markdown H1 | 124 | 2.1% |
| HTML page (soft 404, homepage or app shell) | 929 | 15.7% |
| Redirected to a different path | 22 | 0.4% |
| Empty file | 27 | 0.5% |
| Not found or error | 4,120 | 69.8% |

The HTML row is the trap. Many sites answer every unknown path with HTTP 200 and a web page. A study that counted every HTTP 200 response as adoption would report 30.2% (1,782 sites); the real figure is less than half that.

The headline has a 95% interval of 10.7% to 12.4%. Counting all 10,000 domains, including the sites that did not answer us, gives a floor of 6.8%.

## Findings

### Who publishes llms.txt

| Tranco rank | Live sites | Valid llms.txt | Share |
|---|---|---|---|
| 1 to 1,000 | 534 | 81 | 15.2% |
| 1,001 to 5,000 | 2,323 | 289 | 12.4% |
| 5,001 to 10,000 | 3,045 | 310 | 10.2% |

The gap between the top 1,000 and the bottom tier is 5.0 percentage points, with a 95% interval of 2.0 to 8.5, so it is unlikely to be chance. Split into ten bands of 1,000 sites, the decline is uneven: 15.2% in the top band, 9.1% in the next, between 11.1% and 14.4% from rank 2,001 to 8,000, and 6.2% in the last. A trend test over the ten bands confirms a downward slope (z = −4.24), but rank explains only part of who publishes.

The highest-ranked publishers are mostly infrastructure and developer platforms: cloudflare.com, github.com, fastly.net, digicert.com, wordpress.org, adobe.com, opera.com, sentry.io, samsung.com, wordpress.com, dropbox.com, shopify.com, ubuntu.com, stripe.com and hubspot.com are among them. Documentation-heavy companies, whose customers already ask AI assistants how to use their products, are well represented.

### The HTML answers are catch-all pages, not files

On 28 September we also requested /llms-7f3a9c21b8e4d056.txt from every site, an address no site publishes. It works as a control: a site that returns the same kind of response there is answering any address, not serving a file.

| State at /llms.txt on 26 September | Sites | Control address also returned HTTP 200 |
|---|---|---|
| HTML page | 929 | 897 (96.6%) |
| Valid llms.txt | 680 | 75, none with the same content as the file |

So the HTML responses are what they look like: generic pages served for any unknown address. The valid files are real. 605 of the 680 sites (89.0%) returned an error at the control address, and none of the 75 that answered it returned the same text as their llms.txt.

### Two fetches, two days apart

We fetched /llms.txt from all 5,902 sites again on 28 September and classified the response with the same rule. 99.2% of sites were in the same state (Cohen’s kappa 0.983). Of the 680 valid files, 677 were still valid; 3 had gone and 11 new ones appeared, giving 688 (11.7%) on the second date. One fetch was enough for the headline, and the small net change fits the rapid growth other trackers report.

### What the files contain

- Median size 7,706 bytes; 10% of files are larger than 54,044 bytes.
- Median of 36 links.
- 95.0% organize links under H2 sections and 80.6% include the blockquote summary.
- Only 14.6% link to Markdown (.md) versions of pages, the part of the proposal that saves an AI system the most work.

Read as a ladder, each part of the proposal loses sites:

| Meets | Valid files | Share |
|---|---|---|
| Markdown H1 title | 680 | 100% |
| plus a blockquote summary | 548 | 80.6% |
| plus H2 sections of links | 500 | 73.5% |
| plus links to Markdown pages | 81 | 11.9% |

### llms-full.txt is much rarer

123 live sites (2.1%) serve a valid llms-full.txt. Of the 680 sites with a valid llms.txt, 111 (16.3%) also publish the full-text file.

### Some publishers block the crawlers the file is written for

33 sites (4.9% of llms.txt publishers) also block GPTBot, ClaudeBot or OAI-SearchBot from the whole site in robots.txt. The file invites AI systems in; the robots.txt turns some of them away.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | Tranco top 10,000, 5,902 live sites, September 2026, strict validity rule | 11.5% valid |
| Rankability tracker | Tranco top 1,000, checked September 2026 | 9.3% (93 of 1,000) serve llms.txt or llms-full.txt |
| Casey Burridge (HTTP Archive data) | 7,504 crawled sites from the top 10,000, June 2026 | 5.61% valid |
| SE Ranking | About 300,000 domains, November 2025 | 10.13% |
| HTTP Archive, Web Almanac 2025 | Whole web, July 2025 crawl | 2.13% of desktop sites |

Adoption is rising fast: Burridge’s tracking started at 1.04% of the top 10,000 in July 2025. Differences between the figures also come from validity rules and from platform rollouts: Burridge attributes much of the recent growth to Shopify enabling the file across its platform, and the Web Almanac attributes 39.6% of files to one SEO plugin.

**Does anything read it?** Ahrefs examined llms.txt requests across 137,210 sites it monitors and found that “97% of those files received zero traffic in May 2026”; AI retrieval bots made 1.1% of the requests that did arrive. SE Ranking found no correlation between having the file and being cited by AI. Google has said it does not use llms.txt, and in 2025 Google’s John Mueller said none of the AI services had said they were using it. Anthropic publishes its own llms.txt, which is not the same as its crawlers reading yours.

Sources: [Rankability](https://www.rankability.com/blog/llms-txt-adoption/); [Casey Burridge](https://caseyrb.com/blog/state-of-llms-txt-adoption/); [SE Ranking](https://seranking.com/blog/llms-txt/); [HTTP Archive Web Almanac 2025](https://almanac.httparchive.org/en/2025/seo); [Ahrefs](https://ahrefs.com/blog/llmstxt-study/).

## What this means

This study separates three kinds of claim.

**Established by this data.** About one in nine live top sites publishes a valid llms.txt, more among the most-visited. Counting any response at the address roughly doubles that figure, because most of the extra responses are catch-all pages. The count was stable over two days. Most files follow the basic format; few link to Markdown versions of pages.

**Supported interpretation.** Publication is concentrated among developer platforms and documentation-heavy companies, and platform rollouts move the numbers more than individual decisions do. A file is cheap to publish, which fits wide adoption without evidence of benefit.

**Open questions.** Whether AI systems request the file, whether they use it to reach pages, and whether it changes citation or answers. None of these has been tested with a control: a comparison of sites with and without the file, or of pages listed and not listed in it, under otherwise equal conditions.

### What to do with it

- **Treat llms.txt as a small, cheap convenience, not a ranking factor.** Nothing here shows it changes whether an AI system cites you.
- **If you publish one, make it valid.** An HTML page at /llms.txt is worse than nothing for a tool that expects Markdown. Serve text, start with an H1, add the summary, and link to clean Markdown versions of key pages.
- **Keep it consistent with robots.txt.** A file that welcomes AI systems next to rules that block them sends mixed instructions.

## What we will test next

Each open question above can be turned into a test. These are hypotheses, not findings:

1. **Discovery.** AI crawlers request a valid /llms.txt more often than an equally sized control file at another address on the same site. Test: server logs from sites that publish both.
2. **Retrieval.** Pages linked from llms.txt are fetched by AI user agents more often than similar pages on the same site that are not linked. Test: matched pages, logs over several weeks.
3. **Citation.** Adding a file does not by itself change how often a site is cited. Test: repeated AI answers for the same questions before and after publication, against sites that did not publish.
4. **Platform rollouts.** When a platform enables the file for all its customers, any advantage for an individual site shrinks. Test: track cited Shopify stores across the rollout.

## Methodology

- **Sample:** Tranco list L5PZ4, top 10,000 domains, fetched 26 September 2026; 5,902 had a live homepage (status below 400, more than 200 bytes).
- **Collection:** one HTTPS GET each for /llms.txt and /llms-full.txt, following redirects, with an identifying research user agent.
- **Validity rule:** HTTP 200, not HTML (by content type or markup), final path still /llms.txt, at least 20 characters, first non-blank line a Markdown H1. llms-full.txt uses the same rule without the H1 requirement.
- **Control address and second fetch:** on 28 September 2026 we requested /llms.txt again and /llms-7f3a9c21b8e4d056.txt once from each of the 5,902 live sites, with the same collector and rule. A valid file counts as a catch-all when the control address returns the same text.
- **Statistics:** Wilson 95% intervals for shares, a Newcombe interval for the tier difference, and a Cochran-Armitage test for the trend across ten rank bands.
- **Update schedule:** quarterly.
- **Version 1.1** (28 September 2026) follows an external methodology review. It adds the stage table, the control address, the second fetch, the intervals and the conformance ladder, and separates findings from interpretation. No 1.0 figure changed.

## Limitations

- Sites that blocked our user agent at the homepage are outside the denominator.
- Presence is not use: we measured publication, not whether AI crawlers fetch the file.
- Two fetches, two days apart; files are being added quickly, so the next edition may differ materially.
- One user agent. A site may answer AI crawlers’ user agents differently from ours.
- The rank bands are uneven, so the tier figures describe these sites, not a smooth rule about popularity.

## Data and downloads

- Per-domain results: [s2_llms_by_domain.csv](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain.csv) and [JSON](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain.json)
- Version 1.1 per-domain file with both dates, the control address and the conformance ladder: [s2_llms_by_domain_v11.csv](https://underneath.agency/research-data/llms-txt-adoption-study/s2_llms_by_domain_v11.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/llms-txt-adoption-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/llms-txt-adoption-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *How many websites have an llms.txt file? 2026 adoption data*. Underneath Research. https://underneath.agency/research/llms-txt-adoption-study

## Frequently asked questions

### What percentage of websites have an llms.txt file?

In September 2026, 11.5% of the live websites in the Tranco top 10,000 served a valid llms.txt. Among the top 1,000 sites the figure is 15.2%.

### How do I check whether a site’s llms.txt is valid?

Open /llms.txt in a browser. A valid file is plain text, not a web page, and its first line is a Markdown heading starting with “# ”. If you see the site’s normal design or a “page not found” page, the site does not have one, even if the address loads.

### What is the difference between llms.txt and llms-full.txt?

llms.txt is a short index: a title, a summary and links to the most useful pages. llms-full.txt contains the full text of the site’s key content in one Markdown file. 2.1% of live top sites publish llms-full.txt.

### Does publishing llms.txt get a site cited by AI assistants?

Our data cannot show that. It measures who publishes the file, the first of several stages between a website and an AI answer, not whether AI systems request or use it. Treat llms.txt as a low-cost convenience for AI tools, not as a substitute for crawlable pages and clear facts.

## Related research

- [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)

## Related guides

- [Can our server logs show what content AI bots are looking for on our site?](https://underneath.agency/resources/ai-bot-server-logs-content-demand)
- [How much web content is already optimized for AI search?](https://underneath.agency/resources/how-much-web-content-is-optimized-for-ai-search)
- [The AI search glossary, defined plainly.](https://underneath.agency/resources/glossary)
- [How do developer tool companies win users when developers ask AI first?](https://underneath.agency/resources/developer-tools-ai-search)
- [What GEO practices are actually supported by research?](https://underneath.agency/resources/what-geo-practices-does-research-support)

---

This is the Markdown twin of https://underneath.agency/research/llms-txt-adoption-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
