Guide · AI search

Is tracking ChatGPT alone enough to measure our AI search visibility?

No: for the same question, AI engines cite mostly different sources and often recommend different brands, so one engine shows only part of your visibility. In one four-engine test, ChatGPT captured 42.6% of all the pages the engines cited, the most of any engine but still less than half. Tracking ChatGPT is a sensible start, not a measure of “AI search” as a whole.

The short version

  1. In a June 2026 test of 15 commercial prompts, no single engine captured half of the pages cited across ChatGPT, Google, Perplexity and Copilot; ChatGPT came closest at 42.6%.
  2. In the same test, 96.4% of all cited pages appeared on only one engine.
  3. On shopping questions asked from the Netherlands, ChatGPT and Gemini shared only 5.4% of the websites they displayed for the same question.
  4. In our 80-question study, two assistants’ recommended brands overlapped by only about a third (0.327), and all four named the same first pick for just 10.0% of questions.
  5. Tracking tools that query the developer version of ChatGPT see different sources from the consumer app: the two shared only 12.0% of displayed websites.

How much of the AI source landscape does ChatGPT cover?

Less than half in the one study that measured it directly, and other engines each covered much less.

Tannenbaum (opens in a new tab), founder of a company that sells AI visibility software, ran 15 US commercial prompts about that software category on four engines on 6 June 2026. For the ten prompts all four answered, he pooled every page cited by any engine. ChatGPT cited 42.6% of that pool, Google 24.3%, Perplexity 23.6% and Copilot 11.4%. The pool was 2.43 times as large as the broadest single engine’s list.

The overlap between engines was close to nothing. Of 528 distinct pages cited that day, 96.4% appeared on one engine only. In 84.9% of engine pairs, the two engines cited no page in common for the same prompt. ChatGPT and Google shared no cited page in any of their 11 comparisons. The author flags his conflict of interest and the narrow, self-referential topic.

EngineShare of all pages cited by the four engines
ChatGPT42.6%
Google24.3%
Perplexity23.6%
Copilot11.4%

Do other studies find the same split between engines?

Yes, the pattern of low overlap repeats across independent audits, countries and topics.

Uberti-Bona Marin and colleagues (opens in a new tab), academics with no vendor tie, asked 117 real shopping questions in September 2026 from the Netherlands. For the same question at the same time, ChatGPT and Gemini shared only 5.4% of the websites they displayed. In 76.7% of comparisons they shared none.

Martinez’s survey (opens in a new tab) of the research reports similar gaps elsewhere. One earlier audit found only 26% of domains were cited by both Bing Chat and Perplexity. Another found 53% of domains cited by Google’s AI Overviews did not appear in Google’s own top 10 results. Our guide on whether AI engines cite the same sources gathers these audits.

Google’s ranking does not stand in for ChatGPT either. In our Google rankings study, only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question, against 25.7% for Claude.

Do engines also disagree on which brands they recommend?

Yes: different engines often name different brands for the same buyer question, not just different sources.

In our four-assistant study, we put 80 US buyer questions to ChatGPT, Gemini, Perplexity and Claude in September 2026. Two assistants’ recommended options overlapped by 0.327 on average, about a third. 66.3% of the options recommended for a question came from one assistant only. All four agreed on the first pick for 10.0% of questions.

This is not just run-to-run noise. Each assistant’s answers overlapped with its own repeat runs about twice as much as with another assistant’s. Of the brands one assistant named in at least three of five runs, 50.0% to 56.6% were never named by the other assistant in any of its five. Their taste in page quality is closer, as whether AI engines prefer the same content explains.

Do engines differ in how they show sources and brands?

Yes, so the same metric can mean different things on each engine.

A vendor study by Kumar at Ranqo (opens in a new tab) summarizes a January 2026 test of 50 questions about CRM software on five engines. The share of answers that included source links ranged from 95% on Perplexity to 15% on ChatGPT and 10% on Claude. A “citation share” on ChatGPT is therefore built from far fewer links than the same figure on Perplexity.

Language can create a similar blind spot. Żatuchin (opens in a new tab), who works for a brand-monitoring firm, studied 66 European brands in twelve languages. Asking in a brand’s home language instead of English raised how often local brands were recommended by 0.80 on a 0 to 1 scale, against 0.15 for global brands. A single-engine, English-only tracker would miss both effects.

Does it matter which version of ChatGPT a tool tracks?

Yes: the developer version that many tools query does not reproduce what consumers see in the app.

The Dutch audit paired each consumer-app question with the same question sent to OpenAI’s developer interface a fraction of a second later. The two shared an average of 12.0% of their displayed websites. Ask any tool which version of each engine it queries, from which country, and whether it is logged in.

What should you do about it?

Track the engines your buyers actually use, and report each one separately before combining them.

  1. Find out where your buyers ask. Include Google’s AI features, since they appear on the Google results page itself.
  2. Track at least two or three engines, with the same fixed set of buyer questions on each.
  3. Report every engine on its own. Combine engines only with explicit weights, never by adding raw citation counts.
  4. Check whether a tool queries consumer apps or developer interfaces, and in which country and language.
  5. When one engine moves, check the others before acting. A change on one may not appear anywhere else.

If you want help designing a multi-engine tracker, see our generative engine optimization service.

What does the research not tell us yet?

It does not show how much each engine is worth to a business, or whether one engine predicts the others.

  • The four-engine coverage figures come from 15 prompts in one category on one day, by a vendor.
  • No study here weights engines by how many buyers use them, so a coverage share is not a traffic share.
  • Overlap at the level of brands, not just sources, has been measured on 80 questions or fewer.
  • Claude, Copilot and Grok appear in few studies.
  • Whether visibility on one engine leads to visibility on another over time has not been studied.

Frequently asked questions

Do ChatGPT and Google AI cite the same sources?

Rarely. In one test of 15 prompts, ChatGPT and Google shared no cited page in any of 11 comparisons, and 96.4% of all cited pages appeared on only one engine.

Which AI engine should we track first?

The one your buyers use most, but not only that one. In one four-engine test, the broadest engine still covered just 42.6% of the pages cited across all four.

Do Perplexity and ChatGPT recommend the same brands?

Often not. In our study, two assistants’ recommended options overlapped by about a third (0.327), and all four agreed on the first pick for 10.0% of questions.

Is ranking on Google enough to show up in ChatGPT?

Not reliably. Only 8.3% of the pages ChatGPT cited in our study ranked in Google’s top 10 for the question.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.