---
title: "How often do AI answers say things their sources do not support?"
description: "Often: audits find roughly one in nine to one in three statements in AI search answers are not backed by the sources cited beside them."
canonical: "https://underneath.agency/resources/ai-answers-unsupported-claims"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# How often do AI answers say things their sources do not support?

Often enough to matter: independent audits find that somewhere between one in nine and one in three statements in AI search answers are not backed by the sources listed beside them. The rate depends on the engine, the topic and how strictly “support” is judged. For a brand, it means an answer can describe you inaccurately while still looking well sourced.

## The short version

1. In a 2024 audit of 303 questions, 30.8% of You.com’s statements, 31.6% of Perplexity’s and 23.1% of BingChat’s were not supported by any source the engine listed ([Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349)).
2. In Google’s AI Overviews, 11.0% of 98,020 claims checked over 40 days of US searches were not supported by the pages they cited ([Xu and colleagues](https://arxiv.org/abs/2605.14021)).
3. The main failure is omission, not contradiction: an unsupported claim was about 2.6 times more likely to be missing from the cited pages than contradicted by them (Xu and colleagues).
4. In our own check of AI answers about whether brands are legit, a readable cited page supported the claim fully or in part 72.4% of the time ([our reputation study](https://underneath.agency/research/is-it-legit-ai-reputation-study)).

## How often are AI answers unsupported by their own sources?

In the best-known audit, roughly a quarter to a third of statements had no support in the listed sources. [Narayanan Venkit and colleagues](https://arxiv.org/abs/2410.22349), from Penn State and Salesforce AI Research, ran 303 questions through You.com, BingChat and Perplexity. They split each answer into single statements and had an automated judge check each one against the full text of every listed source.

| Engine (August 2024) | Statements not supported by any listed source | Citations that pointed to a supporting source |
|---|---|---|
| You.com | 30.8% | 68.3% |
| BingChat | 23.1% | 65.8% |
| Perplexity | 31.6% | 49.0% |

Two caveats matter. The automated judge agreed only moderately with human checkers, and these are 2024 versions of products that have changed since. Even so, the authors concluded that none of the three engines reached acceptable performance on most of their measures.

## Are Google’s AI Overviews more faithful?

Yes, by a clear margin, but about one claim in nine was still unsupported. [Xu and colleagues](https://arxiv.org/abs/2605.14021) at Washington University in St. Louis ran 55,393 trending US searches between March 13 and April 21, 2026. They checked every claim in the AI Overviews that appeared (an AI Overview is the AI summary at the top of Google’s results).

Of 98,020 claims, 11.0% were not supported by the pages the Overview cited. Only 4.1% were contradicted by or in conflict with their cited content; a further 7.0% were simply not addressed by any cited page. Most Overviews were largely sound: 41.9% had every claim supported, and only 2.74% had fewer than half supported.

The authors call 11% a ceiling, not a precise rate. Their checker could not read social media and video pages, and if every claim tied to those pages were in fact supported, the rate would fall to about 5.3%. Accuracy also varied by topic, from 94.77% of claims consistent in health to 76.85% in jobs and education, partly because fast-changing facts such as school closings had moved on before pages were checked.

## What do these errors actually look like?

Mostly missing support and misattribution, rather than facts that contradict the source. In the AI Overview data, an unsupported claim was about 2.6 times more likely to be something no cited page mentions than something a cited page contradicts. The answer states a fact, and [the citation beside it does not contain it](https://underneath.agency/resources/ai-citing-pages-for-claims-they-do-not-make).

Misattribution is the second pattern. Venkit’s team found that even when one listed source did support a statement, engines often cited a different one. All 21 experts in their user study flagged answers that cited sources which did not say what the answer claimed, and one called the citations a way for the system “to legitimize itself.”

The citation list can also look broader than the answer really is. [Huang and colleagues](https://arxiv.org/abs/2603.16138) compared 11,000 real search questions across several systems. Google’s AI Overviews cited Reddit and Quora on some questions, yet [drew 22.1% less content from social platforms](https://underneath.agency/resources/do-ai-overviews-downplay-negative-content) than from other sources they cited.

Answers also tend to sound surer than their sources. In the same study, adding web search to an AI assistant [cut hedging language by up to 60%](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions) while keeping confident wording.

## Does citing better sources fix the problem?

No: in the largest audit, how credible a source was had little bearing on whether claims matched it. Xu and colleagues found AI Overviews cited domains that were, on average, more credible than Google’s own first-page results. Yet source quality and claim accuracy were “largely independent,” so the authors argue the problem cannot be fixed by better sourcing alone.

More sources do not guarantee more support either. In the 2024 audit, BingChat listed the most sources, but 36% of the sources it listed [were never cited in the answer text](https://underneath.agency/resources/does-ai-search-use-the-pages-it-cites). The authors warn this can create “a false sense of factual backing from multiple sources.”

## What does this mean for how AI describes a brand?

The same gap shows up in brand questions, and much of it starts with the web’s own inconsistencies. In [our study of “Is this brand legit?” answers](https://underneath.agency/research/is-it-legit-ai-reputation-study) from four assistants about 79 brands, we checked a random sample of 240 cited claims.

Where the cited page could be read, it supported the claim fully or in part 72.4% of the time. Another 25.9% were not supported and 1.7% were contradicted. Many review sites block automated reading, and all coding was done by AI models, so treat these as rough bounds.

The clearest gap concerned how common a problem is. 71.4% of claims saying a complaint was common rested on evidence showing individual reports, or nothing about frequency.

Prices show where many wrong facts come from. In [our pricing study](https://underneath.agency/research/ai-pricing-accuracy-study), 64 quoted software prices differed from the official pricing page. For 39 of them, the same figure was still published on another page of the vendor’s own site.

## What should you do about it?

Treat being cited and being described correctly as two separate goals, and check the second one directly.

1. Ask the main AI assistants the questions your buyers ask about you, and read each claim, not just whether your site is cited.
2. Retire or update old pricing, plan and policy pages. Assistants can repeat any figure you still publish, even on a forgotten help page.
3. State key facts in plain, specific sentences on pages that load without logins or scripts. Since omission is the main failure, a fact that is hard to find invites the answer to fill the gap from somewhere else.
4. Watch the review and forum sites that assistants cite about you, because that is where many negative claims originate.
5. Re-check on a schedule. Every audit here is a snapshot, and the engines change often.

If you want help running these checks across engines, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research measures unsupported claims well, but not how often they are false or how they affect buyers.

- Unsupported is not the same as wrong. Many unsupported claims may be true but drawn from a page the engine did not cite.
- Most audits rely on automated judges. Agreement with human checkers ranged from moderate in the 2024 audit to high in the AI Overview study.
- The big audits cover trending searches, debates and expert questions. Very little independent research looks at product or brand questions specifically.
- Results date quickly. The 2024 audit predates current versions of every engine it tested, and no one has rerun the same test across today’s assistants.
- No study we found measures how an unsupported claim about a company changes what buyers think or do.

## Frequently asked questions

### Are AI search answers accurate?

Mostly, but not reliably. In Google’s AI Overviews, 88.97% of 98,020 checked claims were consistent with the cited pages, but two of three chat-style engines in a 2024 audit left around 30% of their statements unsupported.

### Which AI search engine is most faithful to its sources?

No study ranks today’s engines on the same test. In the 2024 audit, BingChat had the fewest unsupported statements (23.1%) and Perplexity the most (31.6%), but all three products have changed since.

### Can a citation be wrong even when the fact is right?

Yes. Experts in the 2024 user study found true statements attached to sources that never mentioned them, and Perplexity’s citations pointed to a supporting source only 49.0% of the time.

### Why would an AI assistant state something about my company that is not on my website?

Usually because another page says it. In our pricing study, 21 of the 64 prices that differed from the official page also appeared on a third-party page cited for the product.

## Sources

- Narayanan Venkit, Laban, Zhou, Mao and Wu (2024), [Search Engines in an AI Era: The False Promise of Factual and Verifiable Source-Cited Responses](https://arxiv.org/abs/2410.22349), arXiv:2410.22349.
- Xu, Iqbal and Montgomery (2026), [Measuring Google AI Overviews: Activation, Source Quality, Claim Fidelity, and Publisher Impact](https://arxiv.org/abs/2605.14021), arXiv:2605.14021.
- Huang, Goyal, Saha and Chandrasekharan (2026), [Answer Bubbles: Information Exposure in AI-Mediated Search](https://arxiv.org/abs/2603.16138), arXiv:2603.16138.
- Underneath (2026), [“Is this brand legit?” How AI assistants build a reputation](https://underneath.agency/research/is-it-legit-ai-reputation-study)
- Underneath (2026), [How faithfully do AI assistants quote software prices?](https://underneath.agency/research/ai-pricing-accuracy-study)

---

This is the Markdown twin of https://underneath.agency/resources/ai-answers-unsupported-claims. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
