Research · AI assistants

Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?

Search marketing often assumes that ranking on Google is the way into AI answers. That assumption treats an AI answer as one step, when it is several: the assistant decides to search, rewrites the question into its own searches, retrieves candidate pages from an index, and chooses which of them to cite. A page can drop out at any of these stages, and Google rankings only describe one of them.

This study measures how far Google rankings get you through that chain. For 80 US buyer questions on 26 September 2026 we compared every source ChatGPT, Gemini, Perplexity and Claude cited, 2,492 citations in all, with Google’s results for the question as typed. Version 2.0 then asks where the gap opens. We collected Google’s top 100 for the questions and for the 509 searches the assistants ran on them, sorted every citation by the route through which Google could have surfaced it, and modeled which top-ranking pages get cited. It is an observational study: it describes where Google’s view and the assistants’ choices part ways, not what would change them.

The short version

  1. Google’s top 10 for the question explains a minority of citations. 8.3% of the pages ChatGPT cited ranked there (95% interval 5.6% to 11.3%), Gemini 16.6%, Perplexity 14.2% and Claude 25.7%. Google’s own AI Overviews cite top-10 pages 28.7% of the time.
  2. The searches the assistants run explain part of the rest. Adding pages in Google’s top 10 for the assistant’s own searches raises Claude from 25.7% to 35.6% and ChatGPT from 8.3% to 18.5% (for ChatGPT, using searches from a separate run as a stand-in).
  3. A large share of citations are not in Google’s top 100 for the question or for any of those searches: 68.5% for ChatGPT, 66.8% for Perplexity, 49.8% for Gemini and 36.8% for Claude. For ChatGPT, Google rankings describe less than a third of what it cites.
  4. Among the pages Google ranks in its top 10, ranking for the assistant’s own search predicts citation more strongly than ranking for the question. For ChatGPT, once its searches are accounted for, rank for the question no longer predicts citation (odds ratio 0.93, 95% interval 0.55 to 1.57).
  5. Engines differ in ways the questions do not explain. Paired on the same questions, ChatGPT’s citations ranked less often than every other engine’s. Gemini cited ranked pages much less often for questions naming a city (15.6 points fewer, 95% interval 4.2 to 25.1). Version 1.0 also reported that Claude cited ranked pages more often for local questions; with intervals, that difference is not distinguishable from zero and is withdrawn.

Research questions

QuestionAnswered here?
RQ1. How often do cited pages rank on Google for the question as typed?Yes
RQ2. How much of the gap do the assistants’ own searches account for?Exactly for Claude; by proxy for ChatGPT and Gemini; not for Perplexity
RQ3. Among pages Google ranks, what predicts which get cited?Yes, with models clustered by question
RQ4. Do the engines differ on the same questions?Yes
RQ5. Do local questions behave differently?Yes, with small subgroups
RQ6. Would changing a page change whether it is cited?No: nothing was changed
RQ7. Are the results stable across runs, wordings and dates?No: one run on one day

Where Google rankings sit in an AI answer

StageWhat happensWhat this study can see
Search activationThe assistant decides whether to search at allWhether an answer cites anything
Query rewritingThe question becomes one or more searchesClaude’s own searches; stand-ins for ChatGPT and Gemini
RetrievalAn index returns candidate pagesGoogle’s ranking, as a stand-in for an index we cannot see
Selection and citationThe model picks which pages to citeWhich pages were cited, and in what order
Use in the answerThe cited page shapes what the answer saysNot measured

Google rankings are evidence about the retrieval stage only, and only for Google’s index. Search activation already separates the engines: Claude cited sources in 61 of 80 answers and Gemini in 73, while ChatGPT and Perplexity cited sources in all 80. An answer that does not search cannot cite a page however well it ranks.

What we measured

We used the 80 buyer questions of our four-assistant study, ten in each of eight industries, and the sources each assistant cited when answering them on 26 September 2026. ChatGPT and Gemini were queried through their consumer apps, and Perplexity and Claude through their APIs with web search. The primary comparison is with Google’s US top 10 for each question as typed, collected the same day. A citation ranks when its normalized address is one of those results. It shares a site when its domain has some page there.

For version 2.0 we collected Google’s US top 100 on 28 September for the 80 questions and for every distinct search the assistants ran on them, 576 searches in all. For Claude these are the 61 searches recorded in the same answers. For ChatGPT (296 searches) and Gemini (152) they come from a separate API run on the same questions and day, because the consumer apps do not expose their searches; they show the kind of search each assistant runs, not the searches behind these answers. Perplexity’s searches are not available.

Each citation is assigned to the first of four routes that fits, checked in this order.

  • A. In Google’s top 10 for the question as typed.
  • B. In Google’s top 10 for one of the assistant’s own searches.
  • C. Anywhere in Google’s top 100 for the question or one of those searches.
  • D. None of these: Google does not surface the page for any of these searches.

Findings

How often cited pages rank for the question

AssistantCitationsPage in top 10Site in top 10First three
ChatGPT3848.3%22.9%8.3%
Perplexity1,58814.2%26.5%23.8%
Gemini25916.6%23.6%16.3%
Claude26125.7%32.2%22.9%
AI Overviews4,05128.7%43.0%–

“First three” counts only the first three sources of each answer, which removes the effect of how many sources an engine lists. Perplexity lists a median of 20 per answer, the others 3 to 5, and its first three citations rank far more often than the rest (23.8% against 14.2% overall). The AI Overview figures come from 800 Google searches in our citation study, so they are a reference point rather than a like-for-like comparison.

Intervals for the page-level share: ChatGPT 5.6% to 11.3%, Perplexity 11.9% to 16.6%, Gemini 11.3% to 22.2%, Claude 19.4% to 32.0%.

How many answers cite a ranked page

AssistantShare of all 80 answers95% interval
ChatGPT32.5%22.5% to 42.5%
Gemini36.2%26.2% to 47.5%
Claude46.2%35.0% to 57.5%
Perplexity78.8%68.8% to 87.5%

An answer without citations counts as citing none. Version 1.0 divided by answers that had citations, which gives Claude 60.7% and Gemini 39.7%. Perplexity’s high answer-level figure mostly reflects how many sources it lists: at the citation level it is below Gemini and Claude.

Where the gap opens

AssistantABCD
ChatGPT8.3%10.2%13.0%68.5%
Gemini16.6%13.5%20.1%49.8%
Perplexity14.2%n/a19.0%66.8%
Claude25.7%10.0%27.6%36.8%

A is the top 10 for the question, B the top 10 for the assistant’s own search, C the top 100, and D not surfaced by Google. Route B is exact for Claude and a proxy for ChatGPT and Gemini. For Perplexity (n/a), whose searches are unknown, some of its C and D citations would move to B if they were known. Route C is a lower bound: 2.4% of the Google collections returned partial results, and the median collection reached rank 73.5 rather than 100. Route D is the upper bound of what Google cannot explain for these searches.

Three readings follow from the table. For ChatGPT, over two thirds of citations come from pages Google does not surface for the question or for searches of the kind ChatGPT runs, so they must come from another index, another search or a different retrieval process. Claude is the engine most aligned with Google: it has the highest shares in routes A and C and the fewest citations Google does not surface. And the rewritten searches matter, but they are not the main story for any engine: route B adds 10.0 to 13.5 points.

The searches matter less than their existence

We checked whether an engine’s citations match its own searches better than they match the searches another engine ran on the same question. For Claude they do not. 24.5% of its citations are in Google’s top 10 for its own search, 25.3% for the ChatGPT stand-in searches and 29.1% for the Gemini ones. Claude ran one search per answer, and other engines’ searches on the same question find its sources as often. For ChatGPT, its stand-in searches match more of its citations (13.8%) than Gemini’s (9.4%) or Claude’s (4.2%).

What this suggests is that rewriting a question into more specific searches reaches different pages from the question as typed, and that the precise wording of the rewrite matters less than the rewriting itself. Since ChatGPT’s searches here come from a separate run, that part is suggestive, not measured.

Which top-ranking pages get cited

Among the pages in Google’s top 10 for each question, we modeled whether each engine cited them (logistic models with errors clustered by question).

AssistantTop-10 pages citedRanks 1 to 3Ranks 4 to 10
ChatGPT4.7%6.7%3.6%
Gemini6.9%6.8%7.0%
Claude12.8%20.2%8.8%
Perplexity33.0%40.8%28.8%

Higher rank predicts citation for ChatGPT (odds ratio per doubling of rank 0.63), Claude (0.65) and Perplexity (0.74), but not for Gemini (0.97, 95% interval 0.69 to 1.36). Adding whether the page is also in the top 10 for the assistant’s own search changes the picture.

  • Claude: a top-10 page that also ranks for Claude’s own search is more likely to be cited (odds ratio 2.25, 95% interval 1.18 to 4.31), and rank for the question still predicts citation (0.69).
  • ChatGPT, stand-in searches: ranking for the search predicts citation (odds ratio 5.93, 95% interval 2.01 to 17.48), and rank for the question no longer does (0.93).
  • Gemini: neither predicts citation (1.22 for the search, 95% interval 0.55 to 2.72).

Engines compared on the same questions

Differences below are paired by question, so they are not caused by one engine answering easier questions. At the citation level, ChatGPT’s citations ranked less often than Gemini’s (8.3 points fewer, 95% interval 2.9 to 14.0), Perplexity’s (5.9 fewer, 2.8 to 8.8) and Claude’s (17.3 fewer, 10.9 to 24.0). Gemini’s ranked less often than Claude’s (9.1 fewer, 0.4 to 17.5). At the answer level, Perplexity cites a ranked page far more often than every other engine, and the other three are not distinguishable from each other.

National and local questions

Version 2.0 classifies each question by the place it names, listed in the protocol: 19 name a city, 4 a travel destination, 1 a state and 56 none. Version 1.0 used a text rule that counted three questions naming only “the US” as local.

AssistantNationalNaming a cityLocal minus national, 95% interval
ChatGPT8.3%11.1%−6.0 to 7.1 points
Perplexity14.7%15.1%−6.3 to 3.0 points
Gemini21.6%6.9%−25.1 to −4.2 points
Claude20.9%35.7%−1.3 to 25.8 points

Only Gemini shows a difference whose interval excludes zero: for local questions its citations rarely rank. Claude’s higher figure for city questions is within chance for 24 local questions, and it no longer counts as a finding.

Google moves too

Between 26 and 28 September, Google’s top 10 for the same 80 questions changed substantially (median overlap 0.58, mean 0.44, measured as shared pages over all pages). Measured against the 28 September results, ChatGPT’s page-level match falls from 8.3% to 5.2%. A comparison with one day’s rankings is a snapshot of a moving target.

What happened to the Bing comparison

We planned to compare the same citations with Bing’s top 10. For all 80 questions, the Bing results returned by our data provider were unrelated to the question (for example, pages about melamine resin for a question about robo-advisors). A retest with different settings gave the same kind of result. We discarded the Bing data. This says nothing about whether any assistant uses Bing.

Observed, inferred and unknown

What we observe. For 80 buyer questions on one day, most of what the four assistants cited was not in Google’s top 10 for the question. Much of it was not in Google’s top 100 for the question or for searches of the kind the assistants run. The share differs by engine, consistently across the same questions, with ChatGPT furthest from Google and Claude closest.

What we infer. The gap between Google and AI citations opens at several stages. Some of it comes from query rewriting, some from citing pages deeper in the results, and for ChatGPT, Perplexity and Gemini most of it from retrieval that Google’s rankings do not describe. Google rankings are therefore a partial stand-in for AI visibility, and a weaker one for some engines than for others.

What remains unknown. Which index or search each consumer app actually used. Whether the same questions give the same citations on another run, in other wording or on another day. Whether a cited page shaped the answer or was listed without being used. And whether changing a page, its ranking or its structure would change whether it is cited: nothing was changed in this study.

What this means

The points below are our interpretation. They follow from the findings but were not tested.

  • Treat Google rankings as one stage, not the outcome. A page that ranks for the question typed is cited by ChatGPT 4.7% of the time. Ranking helps most where the assistant’s own searches find the same page.
  • Track each engine separately. The same questions produce different overlaps with Google for each engine, and Gemini behaves differently again for local questions.
  • Think in searches, not questions. The assistants turn a buyer’s question into narrower searches, and the pages that rank for those searches are the ones more likely to be cited.
  • Being a trusted site can matter more than ranking one page. For ChatGPT, the site-level overlap (22.9%) was nearly three times the page-level overlap (8.3%): it often cited a different page from a site that ranks.

Where this sits in GEO research

The first generative engine optimization study, by Aggarwal and colleagues (opens in a new tab) (KDD 2024), rewrote pages already placed in the model’s context and measured how visible they became in the answer. That isolates the citation stage: a page that is never retrieved cannot benefit. SAGEO Arena (opens in a new tab) (Kim and colleagues, 2026) tested such rewrites in a full retrieval, reranking and generation pipeline and found that they often hurt retrieval and reranking even where they helped generation. A 2026 critical survey of the field (opens in a new tab) (Martinez) argues that GEO has to be measured stage by stage, with repeated runs and clustered statistics.

This study sits at the boundary between retrieval and citation, on the live consumer systems rather than in a simulated pipeline. It shows that the stages can be told apart in live data: ranking for the question and ranking for the assistant’s own search predict citation separately, and differently for each engine. It does not test any intervention. The question that follows is causal: when the same information is added to a page’s body, its structure or its evidence, does it help retrieval, citation or both, does a gain at one stage cost something at another, and does the answer hold across engines, question types, rewordings and repeated runs? A controlled design that measures each stage separately is needed to answer it. Until then, the figures here describe where the gap is, not how to close it.

Methodology

  • Questions: the 80 US buyer questions of our four-assistant study, ten in each of eight industries; locality classified by the place each question names, listed in the protocol.
  • Citations: sources cited by ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper, US location), Perplexity (sonar) and Claude (Haiku 4.5) through DataForSEO’s LLM Responses API with web search, all 26 September 2026; addresses normalized (tracking parameters, fragments, “www.” and trailing slash removed).
  • Google, question as typed: DataForSEO SERP API, Google.com, United States, English, desktop, top 10 organic results, 26 September 2026.
  • Google, top 100: same settings, depth 100, 28 September 2026, for the 80 questions and every distinct search the assistants ran on them. One search, a “site:” query, returned no results.
  • Assistants’ searches: Claude’s recorded in the same answers; ChatGPT’s and Gemini’s from a separate API run on the same questions and day (proxy); none for Perplexity.
  • Statistics: 95% intervals from a cluster bootstrap that resamples questions (10,000 resamples, seed 23); engine differences paired by question; a difference is reported as a finding only when its interval excludes zero. Candidate models are logistic GEE with robust errors clustered by question.
  • Protocol: written and frozen on 28 September 2026 before the version 2.0 collection; deviations are listed in the methodology file.
  • Bing: collected with DataForSEO’s Bing SERP API and discarded after inspection.
  • Update schedule: quarterly.

Limitations

  • One run per assistant per question, on one day; run-to-run variation is not measured here.
  • The searches for ChatGPT and Gemini come from API models, not the consumer apps whose citations are analyzed, and Perplexity’s are unknown.
  • Google’s top 100 was collected two days after the answers, and Google’s results moved in that time.
  • Google is a stand-in for retrieval; the indexes the assistants actually use are not observed.
  • Local subgroups are small (24 local questions).
  • Observational: no page, ranking or answer was changed, so none of the findings is causal.

What changed in version 2.0

Version 1.0 (26 September 2026) compared citations with Google’s top 10 for the question as typed. Following an external methodological review, version 2.0 adds Google’s top 100 for the questions and for the assistants’ own searches, the four-route breakdown, models of which top-ranking pages are cited, answer-level figures over all 80 questions, 95% intervals clustered by question, paired engine comparisons and an explicit locality list. The citation-level figures are unchanged. Claude’s local-question difference (33.0% against 20.9% in version 1.0) is withdrawn, because its interval includes zero.

Data and downloads

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank? (Version 2.0). Underneath Research. https://underneath.agency/research/ai-citations-google-rankings-study

Frequently asked questions

Does ranking on Google help you get cited by ChatGPT?

Less than people assume. Only 8.3% of the pages ChatGPT cited for 80 buyer questions ranked in Google’s top 10 for the same question, and 68.5% were not in Google’s top 100 for the question or for searches of the kind ChatGPT runs. Among top-10 pages, those that also rank for ChatGPT’s own kind of search are the ones more likely to be cited.

Which AI assistant cites pages that rank on Google most often?

Claude, among the four we tested: 25.7% of its citations ranked in Google’s top 10 for the same question, and 36.8% were not surfaced by Google at all. Google’s own AI Overviews cite top-10 pages 28.7% of the time.

Why don’t AI assistants cite the pages that rank?

Partly because they search with their own rewritten queries, which find different pages. Those searches account for 10.0 to 13.5 points of citations. The larger part of the gap is pages Google does not surface for these searches at all, which points to retrieval that Google’s rankings do not describe.

Does ChatGPT use Bing results?

We could not test it: the Bing results we collected were unrelated to the questions and were discarded. This study says nothing either way.

Is this study proof of what causes AI citations?

No. It is observational: it shows where Google rankings and AI citations diverge, stage by stage, but nothing was changed to test a cause. Whether optimizing a page helps one stage and hurts another needs a controlled experiment.

Free strategy call

Want these numbers for your own category?

We run the same measurements for a business’s own buyer questions and competitors. On a free 30-minute call we’ll take a first look and send you a short written read afterward.