The short version
- In a June 2026 audit of 15 US commercial prompts, ChatGPT cited a mean of 17.8 web addresses per prompt, Google 12.1, Perplexity 10.0 and Copilot 4.5; the author sells AI visibility software.
- In a university audit of 2,848 answers on politics, health and the environment, ChatGPT cited 14.7 sources per answer, Copilot 8.4, Perplexity 7.9 and Gemini 7.3.
- Through developer access the order can reverse: one study counted 36 to 40 citations per answer on Gemini and 5 to 7 on SearchGPT.
- In our own test of 80 US buyer questions, Perplexity listed a median of 20 sources per answer and the other three assistants 3 to 5.
- Engines read far more than they show: ChatGPT’s developer version returned 37.38 pages per question from its searches but cited 3.22.
How many sources does each engine cite?
Most engines cite somewhere between a handful and a few dozen sources, and the counts vary widely by study. The table brings together the largest recent measurements.
| Study | ChatGPT | Google AI | Perplexity | Copilot | Gemini |
|---|---|---|---|---|---|
| Tannenbaum, June 2026, 15 US commercial prompts | 17.8 | 12.1 | 10.0 | 4.5 | not tested |
| Allaham and Diakopoulos, 2026, 712 queries on politics, health, environment | 14.7 | not tested | 7.9 | 8.4 | 7.3 |
| Zhang, He and Yao, 2026, 602 controlled prompts | 6.88 | 12.06 | 16.35 | not tested | see Google |
Tannenbaum (opens in a new tab) ran one day of prompts about AI visibility software in a US setting. The paper discloses that the author founded the company whose tool collected the data, so treat it as a vendor study. Allaham and Diakopoulos (opens in a new tab) used the public web versions of each engine and found ChatGPT cited the most. Zhang, He and Yao (opens in a new tab), three independent researchers, re-analyzed a public dataset and found the opposite order, with Perplexity highest; in their data “Google” means AI Overviews and Gemini together.
Google’s own surfaces differ too. In our AI Overview study of 481 US AI Overviews, the median was 8 citations per AI Overview, the AI summary at the top of Google’s results. On the searches where both appeared, our comparison of AI Mode and AI Overviews found a median of 8 against 4.5 for AI Mode, Google’s separate chat-style search tab.
Why do studies get such different numbers?
Topic, date and the way an engine is reached all change the count. None of the studies above is wrong; they measured different things.
Topic matters. In the Allaham and Diakopoulos audit, answers on the environment carried 10.9 citations on average, against 8.4 for health and 7.7 for politics. Access route matters even more. Sielinski (opens in a new tab), who works for an AI visibility company, queried engines through their developer interfaces over nine days in February 2026, on three consumer product topics. It found median citations of 36 to 40 on Gemini, 19 to 22 on Perplexity and 5 to 7 on SearchGPT. Our guide on tracking AI visibility through developer access explains why that gap matters for monitoring.
The question itself also matters. In the Zhang, He and Yao dataset, harder questions with several conditions pushed ChatGPT down to 3.4 citations, while Google rose to 12.6 and Perplexity to 17.7. Our own test of 80 US buyer questions on 26 September 2026 found yet another pattern, described in our Google rankings study: Perplexity listed a median of 20 sources per answer and ChatGPT, Gemini and Claude 3 to 5.
Do some AI answers cite nothing at all?
Yes, and how often depends heavily on the engine. Some engines answer from memory for a large share of questions.
In a 2025 comparison of six AI search engines with Google and Bing (opens in a new tab), run in incognito browser sessions on trending US topics, the AI engines cited 4.3 web addresses per answer on average, against 10.3 for the classic search engines. 82% of Grok’s answers and 38% of Gemini’s contained no cited website.
Other studies find fewer empty answers. In the politics, health and environment audit, 11.6% of health questions got an answer with no citations. Another 2026 study (opens in a new tab) of 11,000 real search questions found that SearchGPT cited a median of 3 sources, with 97% of answers citing at least one. In our hidden-searches study, none of the 42 answers that ran no web search cited anything.
Do AI engines read more pages than they cite?
Yes, many more: the citations you see are a small slice of what the engine retrieved.
An audit of AI product recommendations (opens in a new tab) could see both layers in ChatGPT’s developer version. Its searches returned a mean of 37.38 pages per question, but the final answer cited 3.22. Only 9.5% of the returned pages made it into the citations.
Our hidden-searches study points the same way from another angle. ChatGPT ran the most searches, yet cited only 2.38 websites per answer, against 6.28 for Gemini and 3.14 for Claude. How many results an assistant keeps matters as much as how much it searches.
Why does the number of citation slots matter for a brand?
Fewer slots mean fewer chances to be cited, so raw citation counts cannot be compared across engines. A share that looks small on one engine may be strong on another.
Tannenbaum puts it simply: a page that is never eligible for a four-source answer faces different odds from one judged in an 18-source answer. The same study found that 96.4% of all cited web addresses appeared on only one of the four engines. The engines are not just citing different amounts; they are citing different pages. Our guide on how concentrated AI citations are across websites looks at who fills those slots.
Counting also distorts totals. Sielinski concluded that raw citation counts are not comparable across platforms and recommended share of citations within each engine instead. If you add up mentions across engines, the engine that cites the most will dominate the total. Our guide on combining engines into one visibility score shows how to weight them instead.
What should you do about it?
Measure each engine separately and judge your share against that engine’s own number of slots. Then:
- Ask your team or agency for citation share per engine, never one blended number across engines.
- Record how many sources each engine cited for your key questions, so a “low” share can be read against a small denominator.
- Track the questions your buyers actually ask; counts change with topic and question complexity.
- Check whether your tracking uses the consumer app or developer access, because the two can differ widely.
- Treat any benchmark as dated; every study here is a snapshot.
If you want help setting up per-engine tracking, see our generative engine optimization service.
What does the research not tell us yet?
The research does not give a stable, current number for any engine. The main gaps:
- Most studies are single snapshots, from one day to a few weeks, and engines change often.
- Topics are narrow: commercial software prompts, politics and health, or consumer products. Other industries are untested.
- Several key figures come from developer interfaces, which do not match what consumers see in the apps.
- No study we reviewed measures how the number of citations affects clicks or sales for a brand.
- The largest four-engine count comes from a vendor with a disclosed conflict of interest and has not been independently repeated.
Frequently asked questions
Which AI search engine cites the most sources?
It depends on the study. ChatGPT cited the most in two 2026 audits of public web versions (17.8 and 14.7 per answer), while Perplexity and Gemini cited more in studies using developer access.
How many sources does a Google AI Overview cite?
About 8 in our data. Our study of 481 US AI Overviews in September 2026 found a median of 8 citations each.
Does ChatGPT cite every page it reads?
No. In one audit its developer version retrieved 37.38 pages per question on average and cited 3.22.
Is a higher citation count better for my brand?
Not by itself. More slots give more chances to appear, but each citation then carries less weight, so compare your share within each engine.
Sources
- Tannenbaum (2026), Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores (opens in a new tab), arXiv:2609.22655.
- Allaham and Diakopoulos (2026), Synthetic Sources?: Auditing Generative Search Engine Citations for Evidence of AI-Generated Sources (opens in a new tab), arXiv:2605.23684.
- Zhang, He and Yao (2026), From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms (opens in a new tab), arXiv:2604.25707.
- Zhang and colleagues (2025), Source Coverage and Citation Bias in LLM-based vs. Traditional Search Engines (opens in a new tab), arXiv:2512.09483.
- Huang and colleagues (2026), Answer Bubbles: Information Exposure in AI-Mediated Search (opens in a new tab), arXiv:2603.16138.
- Sielinski (2026), Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement (opens in a new tab), arXiv:2603.08924.
- Uberti-Bona Marin and colleagues (2026), "If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations (opens in a new tab), arXiv:2609.18729.
- Underneath (2026), Do AI Overviews cite the pages that rank?
- Underneath (2026), AI Mode vs AI Overviews: how different are the sources?
- Underneath (2026), Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?
- Underneath (2026), The hidden searches AI assistants run before they answer