The short version
- 15.2% of the 5,572 top sites with a readable robots.txt block OpenAI’s GPTBot from the whole site (95% interval 14.3% to 16.2%); 14.1% block Anthropic’s ClaudeBot and 16.2% block Common Crawl’s CCBot, the most-blocked AI crawler.
- AI search crawlers are blocked about half as often as training crawlers: 7.4% block OAI-SearchBot (ChatGPT search) and 7.1% block Claude-SearchBot.
- 51.7% of sites that block GPTBot still allow OAI-SearchBot. But 82.7% of those sites never mention OAI-SearchBot: they block GPTBot by name and the search crawler is allowed by default. Only 76 sites, 17.3% of the group, name both crawlers and treat them differently.
- 5.2% of sites block ChatGPT’s search crawler while allowing Googlebot, and most of them (241 of 291) name OAI-SearchBot to do it.
- The top 1,000 sites block most: 30.0% block at least one AI crawler, against 19.2% of sites ranked 5,001 to 10,000. Below the top 1,000 the rate is roughly flat.
- On the same day, ChatGPT cited 123 pages on top sites whose file we could check against the page’s own address, and none of them was closed to OAI-SearchBot. Six were closed to GPTBot, the training crawler, and were cited anyway.
- Perplexity behaved differently: 209 of the 517 pages it cited on those sites (40.4%) were closed to PerplexityBot by the site’s own file, mostly by a rule that names PerplexityBot. Permission stated in robots.txt is not the same thing as exclusion from every AI answer.
Where this study sits
A page has to pass several stages before it shapes an AI answer, and “AI visibility” can mean any of them. We keep them apart because a site can pass one stage and fail the next.
| Stage | The question | In this study |
|---|---|---|
| 1. Permission | Does robots.txt let the crawler in? | Measured, 5,572 files |
| 2. Fetch | Does the crawler request the page? | Not measured |
| 3. Retrieval | Is the page pulled in for a question? | Not measured |
| 4. Citation | Is the page shown as a source? | Cross-checked, one day |
| 5. Use | How much of the answer comes from it? | Not measured |
| 6. Fidelity | Is it represented correctly? | Not measured |
| 7. Outcome | Does a reader click or act? | Not measured |
Stage 1 is the only stage a site controls completely, and the only one this study measures in full. The cross-check at stage 4 uses citations collected for our other studies on the same day. Stage 6 is the subject of our business facts accuracy study, and the relation between Google rankings and citations is covered in AI citations and Google rankings.
What we measured
A robots.txt file tells crawlers which parts of a site they may fetch. Each AI company publishes the name (the “product token”) its crawlers obey, and most now run separate crawlers for separate jobs:
- Training crawlers collect pages to train models: GPTBot (OpenAI), ClaudeBot (Anthropic), Google-Extended (a control token for Gemini training, not a separate crawler), Applebot-Extended, CCBot (Common Crawl), Bytespider (ByteDance), Meta-ExternalAgent, Amazonbot.
- AI search crawlers build the index an assistant searches when it answers with sources: OAI-SearchBot, Claude-SearchBot, PerplexityBot, DuckAssistBot.
- User-triggered fetchers load a page when a person asks the assistant to read it: ChatGPT-User, Claude-User, Perplexity-User, MistralAI-User, meta-externalfetcher.
- Older tokens that many robots.txt files still name: anthropic-ai and Claude-Web (Anthropic) and cohere-ai (Cohere). They count toward the 20 but are not reported separately below.
For every domain we requested /robots.txt once and, for each token, evaluated whether the site root (/) is allowed, using the rules in RFC 9309: the group that names the token applies, otherwise the wildcard (*) group; the longest matching rule wins; allow wins a tie. A site “blocks” a crawler when the root is disallowed for it. Of the 10,000 domains, 5,572 returned a readable plain-text robots.txt; the rest were infrastructure domains without a website, returned an error, or served an HTML page at the address.
Every percentage on this page uses those 5,572 files as the base unless it says otherwise. So 15.2% means 15.2% of top domains with a readable robots.txt, not 15.2% of the top 10,000. The choice of base changes the size of the numbers but not the pattern:
| Base | Domains | GPTBot | OAI-SearchBot | Any AI crawler |
|---|---|---|---|---|
| Readable robots.txt (this page) | 5,572 | 15.2% | 7.4% | 20.7% |
| Readable file or no file | 7,003 | 12.1% | 5.9% | 16.5% |
| Live homepage, readable file or no file | 5,372 | 12.8% | 6.3% | 17.6% |
| All 10,000 domains (lower bound) | 10,000 | 8.5% | 4.2% | 11.5% |
Each column is the share that blocks that crawler from the whole site. A missing file allows everything, so the second and third rows count those domains as not blocking. The last row counts every unreadable file as not blocking, so it is a floor.
Findings
Training crawlers are blocked most
| Crawler | Operator | Purpose | Blocks site | Named |
|---|---|---|---|---|
| CCBot | Common Crawl | Training datasets | 16.2% | 15.2% |
| Bytespider | ByteDance | Training | 15.3% | 13.2% |
| GPTBot | OpenAI | Training | 15.2% | 18.0% |
| ClaudeBot | Anthropic | Training | 14.1% | 15.3% |
| Google-Extended | Gemini training (control token) | 12.3% | 13.7% | |
| Meta-ExternalAgent | Meta | Training | 12.3% | 10.7% |
| Applebot-Extended | Apple | Training (control token) | 11.6% | 10.2% |
| Amazonbot | Amazon | Training and Alexa answers | 11.6% | 10.3% |
Share of 5,572 sites with a readable robots.txt. “Blocks site” means the whole site is blocked, including by a wildcard rule for every unnamed bot; “Named” counts files that mention the token at all, whether to block or allow it.
AI search and user-triggered crawlers are blocked about half as often
| Crawler | Operator | Purpose | Blocks the whole site |
|---|---|---|---|
| PerplexityBot | Perplexity | Search index | 10.9% |
| ChatGPT-User | OpenAI | User-triggered | 9.7% |
| meta-externalfetcher | Meta | User-triggered | 8.3% |
| DuckAssistBot | DuckDuckGo | Search answers | 7.8% |
| Perplexity-User | Perplexity | User-triggered | 7.7% |
| OAI-SearchBot | OpenAI | Search index | 7.4% |
| Claude-User | Anthropic | User-triggered | 7.4% |
| Claude-SearchBot | Anthropic | Search index | 7.1% |
| MistralAI-User | Mistral | User-triggered | 7.1% |
The controls show how unusual AI blocking is: only 2.3% of the same sites block Googlebot, 3.1% block Bingbot and 4.4% block Applebot. GPTBot is blocked 6.6 times as often as Googlebot.
Half the sites that block training still allow AI search, mostly by default
Of the 849 sites that block GPTBot, 439 (51.7%, 95% interval 48.3% to 55.1%) still allow OAI-SearchBot, the crawler OpenAI uses for ChatGPT search. The same split appears at Anthropic: 49.6% of ClaudeBot blockers leave Claude-SearchBot open. Google-Extended works the same way by design: 10.1% of all sites opt out of Gemini training through Google-Extended while still allowing Googlebot.
Whether that split is a choice is a different question, and the files mostly say no. Of the 439 sites, 363 (82.7%) never mention OAI-SearchBot: they block GPTBot by name, and OAI-SearchBot falls through to rules that allow it. Only 76 (17.3%) name OAI-SearchBot and treat it differently from GPTBot. At Anthropic the share is smaller: 48 of 391 (12.3%) name Claude-SearchBot.
| OpenAI crawlers | Sites | Share |
|---|---|---|
| GPTBot allowed, OAI-SearchBot allowed | 4,718 | 84.7% |
| GPTBot blocked, OAI-SearchBot allowed | 439 | 7.9% |
| of which OAI-SearchBot is named in the file | 76 | – |
| GPTBot blocked, OAI-SearchBot blocked | 410 | 7.4% |
| GPTBot allowed, OAI-SearchBot blocked | 5 | 0.1% |
So the pattern is consistent with sites separating training from AI search, but for most of them the separation follows from naming only the training crawler. Robots.txt cannot tell us whether the owners meant it.
A smaller group asks AI search crawlers to stay away but lets Google in
5.2% of sites (95% interval 4.7% to 5.8%) block OAI-SearchBot while allowing Googlebot, and 5.7% block all three AI search crawlers we tested (OAI-SearchBot, Claude-SearchBot and PerplexityBot). This group does look deliberate: 241 of the 291 name OAI-SearchBot in the file. OpenAI’s documentation says sites that opt out of OAI-SearchBot “will not be shown in ChatGPT search answers, though can still appear as navigational links.” We did not test whether these sites appear in ChatGPT or Claude answers.
Blocks that come from a wildcard rule
A crawler that no group names follows the wildcard (*) group. A file written years ago to keep all bots out therefore also blocks crawlers that did not exist when it was written. The rule source differs by crawler type:
| Crawler | Blocked | Named in the rule | Through the wildcard group |
|---|---|---|---|
| GPTBot | 849 | 676 | 173 (20.4%) |
| ClaudeBot | 788 | 601 | 187 (23.7%) |
| PerplexityBot | 607 | 429 | 178 (29.3%) |
| OAI-SearchBot | 415 | 242 | 173 (41.7%) |
| Claude-SearchBot | 398 | 208 | 190 (47.7%) |
Across all AI search crawlers, 39.8% of blocks come through the wildcard group, against 24.9% for training crawlers. 214 files (3.8%) disallow the root for any unnamed bot; 155 of them (72.4%) name no AI crawler at all, and 32 let OAI-SearchBot in by name.
Most files restrict paths, not whole sites
Blocking the root is the strongest setting but not the most common. For OAI-SearchBot, 7.4% of files block the whole site, 3.3% have path rules written for it by name, 68.2% apply general path rules from the wildcard group (for example /admin/ or /search), and 21.0% have no disallow rule that applies to it. Googlebot has a similar profile, apart from the full blocks.
To see whether AI crawlers get less of a site than Google, we tested every path each file itself lists. Among 4,009 sites whose root is open to both OAI-SearchBot and Googlebot, 5.7% close at least one of those paths to OAI-SearchBot but not to Googlebot; for GPTBot it is 6.0% of 3,620 sites. We cannot tell which of a site’s pages matter for AI answers, so this is a count of differences, not of their importance.
The top 1,000 sites block most
| Tranco rank | Readable files | Any AI crawler | 95% interval | GPTBot |
|---|---|---|---|---|
| 1 to 1,000 | 507 | 30.0% | 26.2% to 34.1% | 20.7% |
| 1,001 to 5,000 | 2,196 | 20.5% | 18.9% to 22.3% | 15.4% |
| 5,001 to 10,000 | 2,869 | 19.2% | 17.8% to 20.7% | 14.2% |
Share of each tier’s readable files that block at least one AI crawler, or GPTBot, from the whole site.
The top 1,000 block at least one AI crawler 10.8 percentage points more often than sites ranked 5,001 to 10,000 (95% interval 6.7 to 15.1), and a trend test across rank bands is clear (p < 0.001). The gap survives two checks: among domains with a live homepage it is 27.2% against 18.3%, and without the files that block every unnamed bot it is 24.8% against 16.6%. But the gradient is not smooth. Split into bands of 1,000 ranks, every band from 1,001 to 9,000 falls between 19.1% and 22.9%; the difference sits in the top 1,000 (and a lower last band, 13.1%). Popular domains differ from the rest in type, size and ownership, so this is an association with rank, not an effect of popularity.
Overall, 20.7% of sites block at least one of the 20 AI crawlers from the whole site, and 24.2% mention at least one AI crawler by name.
Content Signals are still rare
Cloudflare’s Content Signals proposal adds a line such as “Content-Signal: search=yes, ai-train=no” to robots.txt. 159 sites (2.9%) carry one. Among them, 99.4% say search=yes, 64.2% say ai-train=no and 73.6% allow ai-input (use of content to ground AI answers). The pattern matches the crawler data: yes to being found, often no to training.
From permission to citation: a same-day cross-check
To see whether stage 1 shows up at stage 4, we matched every source cited in two of our other datasets, both collected on 26 September 2026, against these robots.txt files: the answers ChatGPT, Claude, Gemini and Perplexity gave to 80 buyer questions, and the sources in 481 US AI Overviews. A cited page counts when its domain is in the top 10,000 with a readable file. The engine operator’s own pages (google.com for Google’s answers) are left out. For pages on the domain itself or its www host, we tested the page’s own path and query against the rules for the engine’s search crawler.
| Engine | Search crawler | Cited pages checked | Closed to that crawler |
|---|---|---|---|
| ChatGPT | OAI-SearchBot | 123 | 0 (0.0%) |
| Claude | Claude-SearchBot | 50 | 7 (14.0%) |
| Perplexity | PerplexityBot | 517 | 209 (40.4%) |
| Gemini | Googlebot | 41 | 2 (4.9%) |
| AI Overviews | Googlebot | 979 | 64 (6.5%) |
Pages on subdomains with their own robots.txt are not in the “checked” column. Intervals for the shares are in stats.json.
- ChatGPT matches its published rules. None of the 123 checked pages was closed to OAI-SearchBot (95% interval 0.0% to 3.0%), while 6 were closed to GPTBot. Training opt-outs did not keep pages out of ChatGPT’s answers; search opt-outs, in this sample, did.
- Perplexity cites pages its crawler is told to avoid. The 209 pages come from 37 domains, led by forbes.com (61 pages), nytimes.com (15) and reddit.com (13). In 170 of them the file names PerplexityBot. 194 of the 209 were open to Googlebot and 74 to Perplexity-User. We cannot tell from answers alone whether Perplexity fetched these pages, took them from another search index or relied on text it already held; the files only show that its declared crawler was asked to stay out.
- Claude’s seven are one site. All seven are Yelp search pages whose file names Claude-SearchBot. Claude also cited sites that block ClaudeBot, the training crawler, far more often than their share of top sites would predict (52.7% of its cited pages on top-10,000 domains with a readable file, against 17.8% expected).
- Google’s exceptions are Reddit. All 64 AI Overview citations closed to Googlebot are reddit.com pages; Reddit’s file disallows every crawler, and Google has a data licensing agreement with Reddit announced in 2024. Google-Extended, Google’s training opt-out, made no visible difference: 17.8% of cited pages on top-10,000 domains with a readable file came from sites that use it, against 16.8% expected from their ranks.
Expected rates reweight the share of blocking sites to the rank mix of the cited domains, because AI answers cite popular sites more than the top 10,000 as a whole. This is one day, 80 questions and a few hundred pages per assistant, with citations clustered on a small number of domains, so the results describe that sample, not each engine in general.
How this compares with other studies
Published figures differ mainly because they measure different things. We report them with their definitions rather than as competing answers.
| Source | Sample and date | What was measured | Figure |
|---|---|---|---|
| This study | Tranco top 10,000, 5,572 readable files, September 2026 | GPTBot disallowed at the site root | 15.2% |
| This study | Same | GPTBot named anywhere in robots.txt | 18.0% |
| HTTP Archive, Web Almanac 2025 | Whole web, July 2025 crawl | GPTBot appears in robots.txt | 4.5% of desktop sites |
| Cloudflare | Cloudflare’s top 10,000 domains, 3,816 with robots.txt, June 2025 | Domains disallowing GPTBot | 312 (250 fully, 62 partially) |
| Cloudflare | Cloudflare customers, September 2026 | Dashboard settings, not robots.txt | 17% block training in some way; less than 1% block search bots |
- Top sites block far more than the web as a whole. The Web Almanac counts mentions across millions of mostly small sites; our top-10,000 mention rate is four times higher.
- Training versus search is the same split Cloudflare sees. Its customer settings show training blocked far more often than search; our robots.txt data shows the same direction (15.2% against 7.4% for OpenAI’s two crawlers).
- Content Signals: the only published figure is Cloudflare’s own deployment of the line through its managed robots.txt on over 3.8 million domains, which is not voluntary adoption. Our 2.9% of top sites is, as far as we can find, the first independent count. Google’s John Mueller has said the directive has “no effects whatsoever for any crawler or LLM”.
Sources: HTTP Archive Web Almanac 2025, SEO chapter (opens in a new tab); Cloudflare, June 2025 (opens in a new tab); Cloudflare, September 2026 (opens in a new tab); Cloudflare, Content Signals Policy (opens in a new tab); Search Engine Roundtable (opens in a new tab).
What this means
The following is our interpretation of the numbers, not part of the measurement.
- Blocking training is not the same as leaving AI search. Each AI company runs its training and search crawlers under separate names, so a site can refuse one and allow the other. Most sites in that position got there by naming only the training crawler. A brand that wants to be eligible for AI search should say so explicitly, with a rule for each search and user crawler, rather than rely on defaults that can change.
- Check for accidental blocks. A wildcard rule written years ago for scrapers blocks every new AI crawler too: 3.8% of sites disallow the root for any unnamed bot, and almost three quarters of those files name no AI crawler at all. About two in five blocks on AI search crawlers come through the wildcard group rather than a rule that names them.
- Robots.txt is only half the story. Many sites block AI crawlers at the firewall or CDN instead, which this study cannot see. A brand should test what the crawlers actually receive, not only what the file says.
What this establishes, and what it does not
Each result below supports a narrower claim than the headline might suggest. The last column is the test that would close the gap.
| Result | What it shows | Limit | Next test |
|---|---|---|---|
| 15.2% block GPTBot | What files ask | Not what crawlers do | Server logs |
| Training blocked more than search | Stated preference | Mostly by default | Panel of file changes |
| Wildcard blocks | Old rules reach new bots | Intent unknown | Owner survey |
| Top 1,000 block most | Association with rank | No controls | Matched comparison |
| ChatGPT cites no barred page | Rules and citations agree | One day, 123 pages | Repeated dates |
| Perplexity cites barred pages | Rules do not bind citations | Route unknown | Logs and fetch tests |
Allowing a search crawler does not mean a site is fetched, retrieved or cited, and blocking it does not rule out a citation that arrives through another index or a user-triggered fetch. The files also do not say why sites configure them as they do. Sites that allow AI crawlers differ from sites that block them in size, type and ownership, so any comparison of their AI visibility needs controls for those differences.
Hypotheses for the next edition
These are hypotheses, not findings. The cross-check above bears on the first two; the rest are untested.
- H1. For engines that build answers from their own crawler’s index, permission is close to necessary for citation. For engines that also draw on other indexes, it is not.
- H2. Training opt-outs (GPTBot, ClaudeBot, Google-Extended) do not change whether a site is cited.
- H3. Among sites that allow the search crawler, permission explains little of the variation in citation; relevance and ranking explain more.
- H4. When a site opens a previously blocked search crawler, citations follow only after a delay, and the delay differs by engine.
- H5. Citation and accurate representation can diverge: a page can be cited for claims it does not make.
How the next edition will test them
The robots.txt scan will repeat each quarter on the same 10,000 domains, which turns it into a panel. Sites that change their rules between editions are a natural experiment. We will compare their citations before and after the change with matched sites of similar rank and type that did not change, across the four assistants, a fixed set of questions per industry, repeated runs and at least two dates per edition. A change to an unrelated path serves as a placebo. For the fetch stage, we will use server logs from sites that share them, starting with our own, and test each cited page’s path, not only the root.
Answers vary from run to run, so results will be reported as distributions with intervals, from a model with separate terms for the engine, the question, the domain and the date, rather than as single averages. Cited claims will be checked against the page they cite, so that a rise in citations is not read as a gain when the page is misquoted.
Methodology
- Sample: registrable domains ranked 1 to 10,000 in the Tranco list L5PZ4 (downloaded 26 September 2026), a research ranking that averages several traffic sources.
- Collection: one HTTPS request for /robots.txt per domain on 26 September 2026 (04:16 to 04:38 UTC), following redirects, with a user agent identifying Underneath’s research crawler.
- Inclusion: 5,572 domains that returned HTTP 200 with a plain-text body. Excluded: 2,467 unreachable domains (mostly infrastructure with no website), 1,431 that returned 4xx (under RFC 9309 a missing file means everything is allowed), 485 that served HTML, 45 server errors.
- Evaluation: an RFC 9309 parser (named group, else wildcard group; longest match; allow wins ties), tested against known cases. “Blocked” means the site root is disallowed. A block is “named” when a user-agent line names the token and comes “through the wildcard group” otherwise.
- Uncertainty: 95% Wilson intervals for shares, Newcombe intervals for differences between rank tiers, and a Cochran-Armitage test for trend across rank bands, repeated on domains with a live homepage and without files that block every unnamed bot.
- Crawler purposes: training, search and user-triggered labels are the operators’ own published descriptions as of 26 September 2026. methodology.json records each token with the page its label came from. Operators rename crawlers and change their jobs; each edition re-checks these pages and reports changed labels as changes.
- Citation cross-check (version 1.2): sources cited by ChatGPT, Claude, Gemini and Perplexity for 80 buyer questions and by 481 US AI Overviews, all collected on 26 September 2026 for our other studies. Each cited host is matched to a top-10,000 domain by suffix. Page-level permission is tested with the same parser on the page’s path and query, for pages on the domain or its www host only. Expected rates reweight the population share of blocking sites to the cited domains’ mix of 1,000-rank bands; exact binomial tests are in stats.json.
- Version 1.2 (28 September 2026) adds the pipeline framing, the citation cross-check, the table of what each result establishes, and the hypotheses and design for the next edition. No robots.txt figure changed.
- Version 1.1 (28 September 2026) reanalyzes the same fetch after an external methodology review. It adds the intervals, the alternative bases, the rule-source and path-level counts, and the dated crawler list, and revises the reading of the training-versus-search split. No 1.0 figure changed.
- Update schedule: quarterly; the next edition will report change against this baseline.
Limitations
- Tranco ranks domains, not businesses; the top 10,000 includes content delivery and API domains.
- Robots.txt is a request, not enforcement. Crawlers that ignore it, and sites that block at the network level, are not measured.
- Permission is not crawling, indexing or retrieval; none of those was measured. Citation was cross-checked on one day only.
- The cross-check covers a few hundred cited pages per assistant, concentrated on a few domains, and only pages on top-10,000 domains with a readable file. Citations show that a page was named, not how the engine obtained it.
- The files show what sites ask, not why. Rank differences are associations, with no controls for site type or size.
- “Blocked” is measured at the site root. Path rules are summarized, but we do not know which pages matter for AI answers.
- One fetch per domain on one day. A few servers answer an unfamiliar user agent differently from a browser.
- Results describe popular global sites, not small business websites.
Data and downloads
- Per-domain results for every crawler: s1_robots_by_domain.csv and JSON
- Version 1.1 per-domain file with rule source (named or wildcard) and path restrictions for every crawler: s1_robots_by_domain_v11.csv
- Version 1.2 citation cross-check, one row per cited page with its crawler permission: s1_citation_linkage_v12.csv and the summary linkage_v12.json
- Every statistic on this page: stats.json
- Machine-readable methodology: methodology.json
The data is free to reuse with attribution (CC BY 4.0).
To cite: Underneath. (2026). Which AI crawlers do top websites block? 10,000 sites, 2026. Underneath Research. https://underneath.agency/research/ai-crawler-blocking-study
Frequently asked questions
What percentage of websites block GPTBot?
In our September 2026 scan of the Tranco top 10,000, 15.2% of the 5,572 sites with a readable robots.txt block GPTBot from the whole site. Among the top 1,000 sites the figure is 20.7%.
Does blocking GPTBot keep a site out of ChatGPT search?
Not by itself. GPTBot collects training data. ChatGPT search relies on OAI-SearchBot, and page visits requested by a user come from ChatGPT-User. 51.7% of the sites that block GPTBot still allow OAI-SearchBot, most of them because the file does not mention it. Allowing the crawler makes a site eligible; it does not mean ChatGPT will find or cite it, which this study did not measure.
Can a site be cited by AI if its robots.txt blocks the crawler?
Yes, depending on the engine. In our same-day check, ChatGPT cited none of 123 pages that were closed to OAI-SearchBot, but 40.4% of the pages Perplexity cited on the sites we could check were closed to PerplexityBot, and Google’s AI Overviews cited Reddit pages that Reddit’s file closes to all crawlers. Robots.txt is a request; it is not a guarantee of exclusion.
Which AI crawler is blocked most often?
Common Crawl’s CCBot, blocked by 16.2% of sites, followed by ByteDance’s Bytespider (15.3%) and OpenAI’s GPTBot (15.2%).
What is the difference between Google-Extended and Googlebot?
Googlebot crawls for Google Search, including AI Overviews. Google-Extended is a separate token that only controls whether content is used to train and ground Gemini models. 10.1% of sites opt out through Google-Extended while still allowing Googlebot.
What is a Content-Signal line in robots.txt?
A proposal from Cloudflare that lets a site state, in robots.txt, whether its content may be used for search, as AI input, or for AI training. 2.9% of the sites we checked have one.