---
title: "When AI agents can’t read your site, who tells your story?"
description: "From web searches and other people’s pages. In one test, 42% of an AI agent’s answer came from outside a business whose site it could not read."
canonical: "https://underneath.agency/resources/when-ai-agents-cant-read-your-site"
published: 2026-10-11
updated: 2026-10-11
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# If AI agents can’t read my website, where does their answer about us come from?

When an AI agent cannot read your website, it answers anyway, from web searches and other people’s pages. In a 2026 test across 1,056 real businesses, 42% of the answer about a hard-to-read business came from somewhere other than the business. Those answers were less accurate, less often recommended the business, and more often left out the facts the buyer asked for.

## The short version

1. In 37,927 AI agent tasks about 1,056 real businesses, 42% of the answer came from outside the business when its site was hard for agents to read ([Finder and colleagues](https://arxiv.org/abs/2609.34951)).
2. The agent almost never gives up: in about 99% of tasks that hit a dead end on a site, it answered anyway from what it found elsewhere.
3. Only 7 to 10% of the answer came from what the AI already knew, so the gap is filled by web searches, not memory.
4. Answers built from the business’s own pages got 48.3% of the requested facts right, against 34.3% for answers built from the open web.
5. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study), 15.2% of top websites with a readable robots.txt file block OpenAI’s GPTBot, many through old rules that were not written with AI in mind.

## Where does the answer come from when an agent can’t read your site?

It comes mostly from web search results and third-party pages, not from your site or the AI’s memory. The clearest evidence is a 2026 study by [Finder and colleagues](https://arxiv.org/abs/2609.34951). They ran 37,927 agent journeys, each a buyer question about a business, across four independent harnesses over 1,056 real businesses. A harness here means one AI agent setup: a model plus its search and page-reading tools.

The businesses were split by how easy their sites were for an agent to fetch and read. The two groups were matched on fame, on how well the AI already knew the brand, and on [how widely the brand was mentioned elsewhere](https://underneath.agency/resources/off-site-mentions-vs-site-readability-ai-agents). An AI judge then labeled where each part of every finished answer came from.

| Where the finished answer came from | Easy-to-read site | Hard-to-read site |
|---|---|---|
| The business’s own pages | 78% | 58% |
| Web search results | 12% | 25% |
| Other websites | 3% | 7% |
| What the AI already knew | 7% | 10% |

In the authors’ words, on sites that are not agent-ready, 42% of the answer comes from somewhere other than the business. The share from web search roughly doubled, from 12% to 25%.

## Doesn’t the AI just fall back on what it already knows?

No: memory supplied only a small, steady slice of the answer whether or not the site was readable. Only 7 to 10% of the finished answer came from the model’s training, the same either way. Whatever the site did not provide, more web searching filled in.

The same team measured how this changed over time. On 90 buyer questions put to successive OpenAI releases, the [share of the answer built from memory](https://underneath.agency/resources/do-ai-assistants-answer-from-training-data) fell from 52% in gpt-4.1 to 14% in gpt-5.6. For a business, that means what an agent can read today matters more than what an AI learned about you years ago.

## Does the agent give up when your site blocks it?

Almost never: it searches more and keeps going. In about 99% of tasks that hit a dead end on a site, the agent still answered, built from whatever it found elsewhere.

The harder a site was to read, the more the agent searched. Mean web searches per task rose from 1.8 on the most readable sites to 4.5 on the least readable. Agents were also blocked by the site 2.1 times as often on hard-to-read sites. Each extra search is another chance for the answer to rest on a page you never wrote, never updated, or cannot vouch for.

That matters because third-party pages can be wrong or even planted. [Luo and Chen](https://arxiv.org/abs/2606.13610) tested 12 AI assistants in the lab. A single polluted page yields fooled rates of up to 27%, meaning the assistant recommended a fake product. The risk was higher where the AI knew the real brands less well. That test used saved search results, mostly in Chinese, not the live web.

## Does it change whether the AI recommends you?

Yes: in this study, readable businesses were clearly recommended almost twice as often. Two AI judges from different companies rated each answer, and both had to give the top score. [Agent-ready businesses got a clear recommendation](https://underneath.agency/resources/do-ai-agents-recommend-readable-websites) 20% of the time, against 11% for the others, or 1.9 times as often.

The gap showed up in all four agent setups and all twelve business categories, though the levels differed widely. One setup clearly recommended 5% of the time, another 36%. When agents could not read a site, they hedged. They admitted they could not access the business 4.4 times more often and vouched for it from secondhand sources 3.0 times more often.

## Are answers built from other sources less accurate?

Yes, mainly because they leave out the facts the buyer asked for. Comparing answers about the same business, from the same agent, to the same question, site-built answers got 48.3% of the requested facts right. Answers built from the open web got 34.3%.

[The main failure was omission, not invention](https://underneath.agency/resources/do-ai-agents-invent-or-omit-business-facts). Facts stated wrongly rose only from 4% to 6%, while facts never mentioned grew from 29% to 45%. An answer built from the web was 3.7 times more likely to contain none of the facts the buyer asked for. Accuracy was graded on a subset of 131 businesses, and the gain was largest for pricing questions.

Our own studies point the same way. In [our software pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), most prices that differed from the official pricing page were not invented. For 39 of 64, the same figure appeared on another page of the vendor’s own site, and for 21 more on a third-party page cited for the product. In [our business facts study](https://underneath.agency/research/ai-business-facts-accuracy-study), 42 of 55 phone numbers that differed from the Google profile were on the business’s own website. What a business publishes, everywhere, is what agents repeat.

## How common is it for a site to be hard for agents to read?

Common enough to matter, often by accident. In [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) of the top 10,000 websites, 15.2% of the 5,572 top sites with a readable robots.txt block OpenAI’s GPTBot, its training crawler. Fewer block the crawlers that power AI search: 7.4% block OAI-SearchBot (ChatGPT search) and 7.1% block Claude-SearchBot. Google’s own opt-out raises a related question, covered in [blocking Google-Extended and AI Overviews](https://underneath.agency/resources/does-blocking-google-extended-hurt-ai-overviews).

Many of those blocks are not deliberate choices about AI. Across all AI search crawlers, 39.8% of blocks came from a catch-all rule that applies to every unnamed bot, against 24.9% for training crawlers. A file written years ago to keep bots out also keeps out agents that did not exist then.

Being allowed in is not the same as being easy to read. In [our study of agent-readable websites](https://underneath.agency/research/agent-readable-web-study), 191 of 5,902 live top websites (3.2%) returned a plain-text Markdown version when an agent asked for it. And 54.2% carried no structured data describing the business. Finder and colleagues report that across nearly 100,000 sites, fewer than 1% earned the top grade on their own readiness score.

## Would being readable mean the AI quotes you instead of others?

Not entirely: for broad “best of” questions, AI search leans heavily on third-party sources even when brand sites are open. In a 2025 study of ranking questions such as “best smartphones”, [Chen and colleagues](https://arxiv.org/abs/2509.08919) classified each cited source as brand-owned, earned (reviews and publishers) or social. For well-known brands, ChatGPT’s citations were 93.5% earned and 6.5% brand-owned.

So the two findings answer different questions. When a buyer asks about your business by name, a readable site lets the agent build most of the answer from your pages. When a buyer asks which product is best, reviews and publishers dominate, and your site is one voice among many.

## What should you do about it?

Make sure agents can get into your site, read each page, and find the facts buyers ask about.

1. **Check your robots.txt file for catch-all blocks.** Ask your web team which AI crawlers and user-triggered agents are allowed, and whether any block comes from an old rule nobody meant for AI.
2. **Put key facts in the page text.** Pricing, plans, features and setup steps should be readable without running scripts or opening menus. In the Finder study, the accuracy gain was largest for pricing.
3. **Keep your own facts consistent.** If two pages on your site show different prices or phone numbers, agents may repeat either one, as our pricing and business facts studies found.
4. **Describe your business in structured data.** A basic organization record with links to your real profiles gives agents a clear statement of who you are.
5. **Keep earning third-party coverage.** For “best of” questions, reviews and publishers still carry most of the weight.

If you want help checking what agents see on your site, our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization) covers this.

## What does the research not tell us yet?

The evidence is strong on direction but rests heavily on one study with real limits.

- **The main study is vendor research.** Its authors work for a company that sells an agent-readiness score, and the study uses that company’s own scoring tool to sort sites. The authors say so themselves.
- **AI judges did the labeling.** Where each part of an answer came from, and whether it recommended the business, was judged by AI models, not people.
- **The sample is narrow.** The businesses were mostly software and commerce, mostly English-language, and the questions were business-fact lookups about pricing, features and setup.
- **Content depth was not matched.** Sites that block agents may also publish less, which could widen the gap on its own. Only the accuracy comparison, which compares answers about the same business, is protected from this.
- **No one has yet fixed a site and measured before and after.** The authors name that as the test that would settle cause and effect.

## Frequently asked questions

### Will ChatGPT still talk about my company if it can’t access my website?

Yes. In one large test, agents answered in about 99% of tasks that hit a dead end on a site, using web searches and third-party pages instead.

### Does blocking GPTBot stop my site appearing in ChatGPT answers?

Not by itself. GPTBot is OpenAI’s training crawler. On one day, [our crawler study](https://underneath.agency/research/ai-crawler-blocking-study) found ChatGPT cited six pages closed to GPTBot, while none of its 123 checked citations was closed to its search crawler.

### Do AI agents make up facts about businesses they can’t read?

Mostly they leave facts out rather than invent them. In the Finder study, facts stated wrongly rose only from 4% to 6%, while facts never mentioned rose from 29% to 45%.

### Is being mentioned on other sites enough for AI visibility?

Not for questions about your own business. In the Finder study, how widely a brand was mentioned elsewhere had no confirmed link to whether answers used its pages or recommended it, once readability was taken into account.

## Sources

- Finder, Elovic, Shalev and Yosef (2026), [AX is the New AEO](https://arxiv.org/abs/2609.34951), arXiv:2609.34951.
- Chen, Wang, Chen and Koudas (2025), [Generative Engine Optimization: How to Dominate AI Search](https://arxiv.org/abs/2509.08919), arXiv:2509.08919.
- Luo and Chen (2026), [One Polluted Page Is Enough: Evaluating Web Content Pollution in LLM Recommenders](https://arxiv.org/abs/2606.13610), arXiv:2606.13610.
- Underneath (2026), [Which AI crawlers do top websites block?](https://underneath.agency/research/ai-crawler-blocking-study)
- Underneath (2026), [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)
- Underneath (2026), [Do AI answers match a business’s Google profile?](https://underneath.agency/research/ai-business-facts-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/when-ai-agents-cant-read-your-site. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
