This study measures how far Google rankings get you through that chain. For 80 US buyer questions on 26 September 2026 we compared every source ChatGPT, Gemini, Perplexity and Claude cited, 2,492 citations in all, with Google’s results for the question as typed. Version 2.0 then asks where the gap opens. We collected Google’s top 100 for the questions and for the 509 searches the assistants ran on them, sorted every citation by the route through which Google could have surfaced it, and modeled which top-ranking pages get cited. It is an observational study: it describes where Google’s view and the assistants’ choices part ways, not what would change them.
The short version
- Google’s top 10 for the question explains a minority of citations. 8.3% of the pages ChatGPT cited ranked there (95% interval 5.6% to 11.3%), Gemini 16.6%, Perplexity 14.2% and Claude 25.7%. Google’s own AI Overviews cite top-10 pages 28.7% of the time.
- The searches the assistants run explain part of the rest. Adding pages in Google’s top 10 for the assistant’s own searches raises Claude from 25.7% to 35.6% and ChatGPT from 8.3% to 18.5% (for ChatGPT, using searches from a separate run as a stand-in).
- A large share of citations are not in Google’s top 100 for the question or for any of those searches: 68.5% for ChatGPT, 66.8% for Perplexity, 49.8% for Gemini and 36.8% for Claude. For ChatGPT, Google rankings describe less than a third of what it cites.
- Among the pages Google ranks in its top 10, ranking for the assistant’s own search predicts citation more strongly than ranking for the question. For ChatGPT, once its searches are accounted for, rank for the question no longer predicts citation (odds ratio 0.93, 95% interval 0.55 to 1.57).
- Engines differ in ways the questions do not explain. Paired on the same questions, ChatGPT’s citations ranked less often than every other engine’s. Gemini cited ranked pages much less often for questions naming a city (15.6 points fewer, 95% interval 4.2 to 25.1). Version 1.0 also reported that Claude cited ranked pages more often for local questions; with intervals, that difference is not distinguishable from zero and is withdrawn.
Research questions
| Question | Answered here? |
|---|---|
| RQ1. How often do cited pages rank on Google for the question as typed? | Yes |
| RQ2. How much of the gap do the assistants’ own searches account for? | Exactly for Claude; by proxy for ChatGPT and Gemini; not for Perplexity |
| RQ3. Among pages Google ranks, what predicts which get cited? | Yes, with models clustered by question |
| RQ4. Do the engines differ on the same questions? | Yes |
| RQ5. Do local questions behave differently? | Yes, with small subgroups |
| RQ6. Would changing a page change whether it is cited? | No: nothing was changed |
| RQ7. Are the results stable across runs, wordings and dates? | No: one run on one day |
Where Google rankings sit in an AI answer
| Stage | What happens | What this study can see |
|---|---|---|
| Search activation | The assistant decides whether to search at all | Whether an answer cites anything |
| Query rewriting | The question becomes one or more searches | Claude’s own searches; stand-ins for ChatGPT and Gemini |
| Retrieval | An index returns candidate pages | Google’s ranking, as a stand-in for an index we cannot see |
| Selection and citation | The model picks which pages to cite | Which pages were cited, and in what order |
| Use in the answer | The cited page shapes what the answer says | Not measured |
Google rankings are evidence about the retrieval stage only, and only for Google’s index. Search activation already separates the engines: Claude cited sources in 61 of 80 answers and Gemini in 73, while ChatGPT and Perplexity cited sources in all 80. An answer that does not search cannot cite a page however well it ranks.
What we measured
We used the 80 buyer questions of our four-assistant study, ten in each of eight industries, and the sources each assistant cited when answering them on 26 September 2026. ChatGPT and Gemini were queried through their consumer apps, and Perplexity and Claude through their APIs with web search. The primary comparison is with Google’s US top 10 for each question as typed, collected the same day. A citation ranks when its normalized address is one of those results. It shares a site when its domain has some page there.
For version 2.0 we collected Google’s US top 100 on 28 September for the 80 questions and for every distinct search the assistants ran on them, 576 searches in all. For Claude these are the 61 searches recorded in the same answers. For ChatGPT (296 searches) and Gemini (152) they come from a separate API run on the same questions and day, because the consumer apps do not expose their searches; they show the kind of search each assistant runs, not the searches behind these answers. Perplexity’s searches are not available.
Each citation is assigned to the first of four routes that fits, checked in this order.
- A. In Google’s top 10 for the question as typed.
- B. In Google’s top 10 for one of the assistant’s own searches.
- C. Anywhere in Google’s top 100 for the question or one of those searches.
- D. None of these: Google does not surface the page for any of these searches.
Findings
How often cited pages rank for the question
| Assistant | Citations | Page in top 10 | Site in top 10 | First three |
|---|---|---|---|---|
| ChatGPT | 384 | 8.3% | 22.9% | 8.3% |
| Perplexity | 1,588 | 14.2% | 26.5% | 23.8% |
| Gemini | 259 | 16.6% | 23.6% | 16.3% |
| Claude | 261 | 25.7% | 32.2% | 22.9% |
| AI Overviews | 4,051 | 28.7% | 43.0% | – |
“First three” counts only the first three sources of each answer, which removes the effect of how many sources an engine lists. Perplexity lists a median of 20 per answer, the others 3 to 5, and its first three citations rank far more often than the rest (23.8% against 14.2% overall). The AI Overview figures come from 800 Google searches in our citation study, so they are a reference point rather than a like-for-like comparison.
Intervals for the page-level share: ChatGPT 5.6% to 11.3%, Perplexity 11.9% to 16.6%, Gemini 11.3% to 22.2%, Claude 19.4% to 32.0%.
How many answers cite a ranked page
| Assistant | Share of all 80 answers | 95% interval |
|---|---|---|
| ChatGPT | 32.5% | 22.5% to 42.5% |
| Gemini | 36.2% | 26.2% to 47.5% |
| Claude | 46.2% | 35.0% to 57.5% |
| Perplexity | 78.8% | 68.8% to 87.5% |
An answer without citations counts as citing none. Version 1.0 divided by answers that had citations, which gives Claude 60.7% and Gemini 39.7%. Perplexity’s high answer-level figure mostly reflects how many sources it lists: at the citation level it is below Gemini and Claude.
Where the gap opens
| Assistant | A | B | C | D |
|---|---|---|---|---|
| ChatGPT | 8.3% | 10.2% | 13.0% | 68.5% |
| Gemini | 16.6% | 13.5% | 20.1% | 49.8% |
| Perplexity | 14.2% | n/a | 19.0% | 66.8% |
| Claude | 25.7% | 10.0% | 27.6% | 36.8% |
A is the top 10 for the question, B the top 10 for the assistant’s own search, C the top 100, and D not surfaced by Google. Route B is exact for Claude and a proxy for ChatGPT and Gemini. For Perplexity (n/a), whose searches are unknown, some of its C and D citations would move to B if they were known. Route C is a lower bound: 2.4% of the Google collections returned partial results, and the median collection reached rank 73.5 rather than 100. Route D is the upper bound of what Google cannot explain for these searches.
Three readings follow from the table. For ChatGPT, over two thirds of citations come from pages Google does not surface for the question or for searches of the kind ChatGPT runs, so they must come from another index, another search or a different retrieval process. Claude is the engine most aligned with Google: it has the highest shares in routes A and C and the fewest citations Google does not surface. And the rewritten searches matter, but they are not the main story for any engine: route B adds 10.0 to 13.5 points.
The searches matter less than their existence
We checked whether an engine’s citations match its own searches better than they match the searches another engine ran on the same question. For Claude they do not. 24.5% of its citations are in Google’s top 10 for its own search, 25.3% for the ChatGPT stand-in searches and 29.1% for the Gemini ones. Claude ran one search per answer, and other engines’ searches on the same question find its sources as often. For ChatGPT, its stand-in searches match more of its citations (13.8%) than Gemini’s (9.4%) or Claude’s (4.2%).
What this suggests is that rewriting a question into more specific searches reaches different pages from the question as typed, and that the precise wording of the rewrite matters less than the rewriting itself. Since ChatGPT’s searches here come from a separate run, that part is suggestive, not measured.
Which top-ranking pages get cited
Among the pages in Google’s top 10 for each question, we modeled whether each engine cited them (logistic models with errors clustered by question).
| Assistant | Top-10 pages cited | Ranks 1 to 3 | Ranks 4 to 10 |
|---|---|---|---|
| ChatGPT | 4.7% | 6.7% | 3.6% |
| Gemini | 6.9% | 6.8% | 7.0% |
| Claude | 12.8% | 20.2% | 8.8% |
| Perplexity | 33.0% | 40.8% | 28.8% |
Higher rank predicts citation for ChatGPT (odds ratio per doubling of rank 0.63), Claude (0.65) and Perplexity (0.74), but not for Gemini (0.97, 95% interval 0.69 to 1.36). Adding whether the page is also in the top 10 for the assistant’s own search changes the picture.
- Claude: a top-10 page that also ranks for Claude’s own search is more likely to be cited (odds ratio 2.25, 95% interval 1.18 to 4.31), and rank for the question still predicts citation (0.69).
- ChatGPT, stand-in searches: ranking for the search predicts citation (odds ratio 5.93, 95% interval 2.01 to 17.48), and rank for the question no longer does (0.93).
- Gemini: neither predicts citation (1.22 for the search, 95% interval 0.55 to 2.72).
Engines compared on the same questions
Differences below are paired by question, so they are not caused by one engine answering easier questions. At the citation level, ChatGPT’s citations ranked less often than Gemini’s (8.3 points fewer, 95% interval 2.9 to 14.0), Perplexity’s (5.9 fewer, 2.8 to 8.8) and Claude’s (17.3 fewer, 10.9 to 24.0). Gemini’s ranked less often than Claude’s (9.1 fewer, 0.4 to 17.5). At the answer level, Perplexity cites a ranked page far more often than every other engine, and the other three are not distinguishable from each other.
National and local questions
Version 2.0 classifies each question by the place it names, listed in the protocol: 19 name a city, 4 a travel destination, 1 a state and 56 none. Version 1.0 used a text rule that counted three questions naming only “the US” as local.
| Assistant | National | Naming a city | Local minus national, 95% interval |
|---|---|---|---|
| ChatGPT | 8.3% | 11.1% | −6.0 to 7.1 points |
| Perplexity | 14.7% | 15.1% | −6.3 to 3.0 points |
| Gemini | 21.6% | 6.9% | −25.1 to −4.2 points |
| Claude | 20.9% | 35.7% | −1.3 to 25.8 points |
Only Gemini shows a difference whose interval excludes zero: for local questions its citations rarely rank. Claude’s higher figure for city questions is within chance for 24 local questions, and it no longer counts as a finding.
Google moves too
Between 26 and 28 September, Google’s top 10 for the same 80 questions changed substantially (median overlap 0.58, mean 0.44, measured as shared pages over all pages). Measured against the 28 September results, ChatGPT’s page-level match falls from 8.3% to 5.2%. A comparison with one day’s rankings is a snapshot of a moving target.
What happened to the Bing comparison
We planned to compare the same citations with Bing’s top 10. For all 80 questions, the Bing results returned by our data provider were unrelated to the question (for example, pages about melamine resin for a question about robo-advisors). A retest with different settings gave the same kind of result. We discarded the Bing data. This says nothing about whether any assistant uses Bing.
Observed, inferred and unknown
What we observe. For 80 buyer questions on one day, most of what the four assistants cited was not in Google’s top 10 for the question. Much of it was not in Google’s top 100 for the question or for searches of the kind the assistants run. The share differs by engine, consistently across the same questions, with ChatGPT furthest from Google and Claude closest.
What we infer. The gap between Google and AI citations opens at several stages. Some of it comes from query rewriting, some from citing pages deeper in the results, and for ChatGPT, Perplexity and Gemini most of it from retrieval that Google’s rankings do not describe. Google rankings are therefore a partial stand-in for AI visibility, and a weaker one for some engines than for others.
What remains unknown. Which index or search each consumer app actually used. Whether the same questions give the same citations on another run, in other wording or on another day. Whether a cited page shaped the answer or was listed without being used. And whether changing a page, its ranking or its structure would change whether it is cited: nothing was changed in this study.
What this means
The points below are our interpretation. They follow from the findings but were not tested.
- Treat Google rankings as one stage, not the outcome. A page that ranks for the question typed is cited by ChatGPT 4.7% of the time. Ranking helps most where the assistant’s own searches find the same page.
- Track each engine separately. The same questions produce different overlaps with Google for each engine, and Gemini behaves differently again for local questions.
- Think in searches, not questions. The assistants turn a buyer’s question into narrower searches, and the pages that rank for those searches are the ones more likely to be cited.
- Being a trusted site can matter more than ranking one page. For ChatGPT, the site-level overlap (22.9%) was nearly three times the page-level overlap (8.3%): it often cited a different page from a site that ranks.
Where this sits in GEO research
The first generative engine optimization study, by Aggarwal and colleagues (opens in a new tab) (KDD 2024), rewrote pages already placed in the model’s context and measured how visible they became in the answer. That isolates the citation stage: a page that is never retrieved cannot benefit. SAGEO Arena (opens in a new tab) (Kim and colleagues, 2026) tested such rewrites in a full retrieval, reranking and generation pipeline and found that they often hurt retrieval and reranking even where they helped generation. A 2026 critical survey of the field (opens in a new tab) (Martinez) argues that GEO has to be measured stage by stage, with repeated runs and clustered statistics.
This study sits at the boundary between retrieval and citation, on the live consumer systems rather than in a simulated pipeline. It shows that the stages can be told apart in live data: ranking for the question and ranking for the assistant’s own search predict citation separately, and differently for each engine. It does not test any intervention. The question that follows is causal: when the same information is added to a page’s body, its structure or its evidence, does it help retrieval, citation or both, does a gain at one stage cost something at another, and does the answer hold across engines, question types, rewordings and repeated runs? A controlled design that measures each stage separately is needed to answer it. Until then, the figures here describe where the gap is, not how to close it.
Methodology
- Questions: the 80 US buyer questions of our four-assistant study, ten in each of eight industries; locality classified by the place each question names, listed in the protocol.
- Citations: sources cited by ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper, US location), Perplexity (sonar) and Claude (Haiku 4.5) through DataForSEO’s LLM Responses API with web search, all 26 September 2026; addresses normalized (tracking parameters, fragments, “www.” and trailing slash removed).
- Google, question as typed: DataForSEO SERP API, Google.com, United States, English, desktop, top 10 organic results, 26 September 2026.
- Google, top 100: same settings, depth 100, 28 September 2026, for the 80 questions and every distinct search the assistants ran on them. One search, a “site:” query, returned no results.
- Assistants’ searches: Claude’s recorded in the same answers; ChatGPT’s and Gemini’s from a separate API run on the same questions and day (proxy); none for Perplexity.
- Statistics: 95% intervals from a cluster bootstrap that resamples questions (10,000 resamples, seed 23); engine differences paired by question; a difference is reported as a finding only when its interval excludes zero. Candidate models are logistic GEE with robust errors clustered by question.
- Protocol: written and frozen on 28 September 2026 before the version 2.0 collection; deviations are listed in the methodology file.
- Bing: collected with DataForSEO’s Bing SERP API and discarded after inspection.
- Update schedule: quarterly.
Limitations
- One run per assistant per question, on one day; run-to-run variation is not measured here.
- The searches for ChatGPT and Gemini come from API models, not the consumer apps whose citations are analyzed, and Perplexity’s are unknown.
- Google’s top 100 was collected two days after the answers, and Google’s results moved in that time.
- Google is a stand-in for retrieval; the indexes the assistants actually use are not observed.
- Local subgroups are small (24 local questions).
- Observational: no page, ranking or answer was changed, so none of the findings is causal.
What changed in version 2.0
Version 1.0 (26 September 2026) compared citations with Google’s top 10 for the question as typed. Following an external methodological review, version 2.0 adds Google’s top 100 for the questions and for the assistants’ own searches, the four-route breakdown, models of which top-ranking pages are cited, answer-level figures over all 80 questions, 95% intervals clustered by question, paired engine comparisons and an explicit locality list. The citation-level figures are unchanged. Claude’s local-question difference (33.0% against 20.9% in version 1.0) is withdrawn, because its interval includes zero.
Data and downloads
- Every citation with its route, Google ranks and searches: s23b_citations_pathways.csv and JSON
- Every Google top-10 page and whether each engine cited it: s23b_candidates.csv and JSON
- Every statistic on this page with its interval: stats.json
- Machine-readable methodology and deviations: methodology.json
- Version 1.0 citations and statistics: s23_citations_vs_google.csv and stats_v1.0.json
The data is free to reuse with attribution (CC BY 4.0).
To cite: Underneath. (2026). Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank? (Version 2.0). Underneath Research. https://underneath.agency/research/ai-citations-google-rankings-study
Frequently asked questions
Does ranking on Google help you get cited by ChatGPT?
Less than people assume. Only 8.3% of the pages ChatGPT cited for 80 buyer questions ranked in Google’s top 10 for the same question, and 68.5% were not in Google’s top 100 for the question or for searches of the kind ChatGPT runs. Among top-10 pages, those that also rank for ChatGPT’s own kind of search are the ones more likely to be cited.
Which AI assistant cites pages that rank on Google most often?
Claude, among the four we tested: 25.7% of its citations ranked in Google’s top 10 for the same question, and 36.8% were not surfaced by Google at all. Google’s own AI Overviews cite top-10 pages 28.7% of the time.
Why don’t AI assistants cite the pages that rank?
Partly because they search with their own rewritten queries, which find different pages. Those searches account for 10.0 to 13.5 points of citations. The larger part of the gap is pages Google does not surface for these searches at all, which points to retrieval that Google’s rankings do not describe.
Does ChatGPT use Bing results?
We could not test it: the Bing results we collected were unrelated to the questions and were discarded. This study says nothing either way.
Is this study proof of what causes AI citations?
No. It is observational: it shows where Google rankings and AI citations diverge, stage by stage, but nothing was changed to test a cause. Whether optimizing a page helps one stage and hurts another needs a controlled experiment.