The short version
- Gemini produced between 19.9 and 50.1 citations per answer depending on the topic, against 4.2 to 6.3 for ChatGPT search, so pooled raw counts are dominated by Gemini (Sielinski, 2026 (opens in a new tab)).
- In a test of 15 commercial prompts across four AI engines, 96.4% of cited pages appeared on only one engine (Tannenbaum, 2026 (opens in a new tab)).
- Across 80 buyer questions, ChatGPT, Gemini, Perplexity and Claude shared about a third of their recommended options and agreed on the top pick for 10.0% of questions (our brand agreement study).
- For one brand over five weeks, Perplexity’s mention rate climbed from 37% to 62% while ChatGPT’s slid from 45% to 20%, a split a single score would flatten (Kumar, 2026 (opens in a new tab)).
Why not just add up results across all the engines?
Because the engines produce very different volumes, so the biggest one swamps the rest. Sielinski (opens in a new tab) measured citations across ten topics. Gemini produced between 19.9 and 50.1 citations per answer, Perplexity between 14.5 and 42.2, and ChatGPT search between 4.2 and 6.3.
His advice is direct. Adding raw citation counts across platforms weights each one by its volume, so a dense platform overshadows sparse ones. The fix is to turn each engine’s results into a share first, then combine the shares, so each engine counts equally regardless of volume. The same applies to the error ranges: work them out per engine, then combine.
Volume differences also show up in other studies. In a June 2026 test, ChatGPT cited 17.8 distinct pages per prompt on average and Microsoft Copilot 4.5 (Tannenbaum, 2026 (opens in a new tab)). A page competing for a slot in a four-source answer faces a different contest from one in an 18-source answer. Our guide to how many sources each engine cites compares the counts across studies.
Do the engines even measure the same thing?
Not really; they mostly cite different pages and often recommend different brands. Tannenbaum (opens in a new tab) put 15 commercial prompts to ChatGPT, Copilot, Google and Perplexity on 6 June 2026. Of the 528 distinct pages cited, 96.4% appeared on only one engine. In 84.9% of engine pairs, the two engines shared no cited page at all for the same prompt. Other studies of whether AI engines cite the same sources point the same way.
No single engine was a good stand-in for the others. The broadest one, ChatGPT, captured 42.6% of all the pages the four engines cited between them; Copilot captured 11.4%. The author is affiliated with Aiso Boost Ltd., and the prompts were about AI visibility software, an unusual topic, so the exact figures may not transfer. Still, it is one reason tracking ChatGPT alone falls short.
Recommendations diverge too. In our brand agreement study, two assistants’ recommended options for the same question overlapped by 0.327 on a scale where 1 means identical. A Dutch audit of product questions (opens in a new tab) found the ChatGPT and Gemini apps shared only 5.4% of their cited websites for the same question.
What can a single combined score hide?
It can hide engines moving in opposite directions. Ranqo, which sells AI visibility tracking, followed one brand, ParallelDots, for five weeks on the same prompts (Kumar, 2026 (opens in a new tab)). Perplexity’s mention rate climbed from 37% to 62% while ChatGPT’s slid from 45% to 20%. An average of the two would look almost flat while both engines changed sharply.
The author treats this as one illustrative brand, not proof. But it shows the risk: a blended number can say “stable” when the right response is to work on one engine and protect gains on another.
A combined score can also hide differences between models from the same company. Kato and colleagues (opens in a new tab) tracked one brand, Glasp, across 2,240 answers from two OpenAI models. Overall, the newer model named it in 378 of 1,120 answers and the older one in 311 of 1,120. Yet the older model led on the English questions, and rates varied widely by use case. The authors conclude that one overall rate is inadequate.
If we do combine engines, how should we weight them?
By a stated, defensible rule, such as each engine’s share of your buyers, and never by raw volume. Any weighting is a choice that changes the answer. Ranqo’s own cross-engine figure, for example, weights ChatGPT at 0.30, Gemini, Perplexity and Claude at 0.20 each and Grok at 0.10.
Martinez (opens in a new tab), in a survey of GEO measurement, argues that a single score conceals the choices behind it when several weightings are reasonable. His advice is to report the range of scores across those weightings. When comparing brands or periods, apply the same weights to both, rather than comparing scores built on different mixes.
In practice, that means three numbers per engine and a clearly labeled blend. If your buyers mostly use ChatGPT, weight it heavily and say so. If you do not know where your buyers are, show the blend under two or three plausible weightings.
Do different engines need different measurement plans?
Yes, because each engine varies in its own way. Sielinski (opens in a new tab) found each engine had a typical level of repeat consistency regardless of answer length: repeated runs of a question shared about 0.30 of their sources on Gemini, 0.40 on ChatGPT search and 0.50 on Perplexity.
Engines also differ in how concentrated their sources are. A Swiss study scored citation concentration at 0.782 for Google AI Mode and 0.671 for Perplexity, on a scale where 1 means one site takes everything (Schulte, Bleeker and Kaufmann, 2026 (opens in a new tab)). Its advice is to set a separate baseline for each engine rather than one threshold for all.
They differ in whether they show links at all. In Ranqo’s study of CRM software questions, Perplexity included source links in 95% of answers and Claude in 10%. A citation-based score is therefore partly a measure of each product’s design, not just of your brand.
What should you do about it?
Keep each engine as its own line, and treat any combined score as a summary on top. Steps:
- Measure each engine separately, with enough repeated runs for that engine, and report its own rate and range.
- Convert each engine’s results to shares before combining, so a high-volume engine does not dominate.
- Choose and publish the weights, ideally based on where your buyers actually ask. Show how the blend changes under other reasonable weights.
- Watch for engines moving in opposite directions. A flat blend with diverging engines needs action, not reassurance.
- Keep mentions, citations and first-place picks as separate measures; they behave differently on each engine.
- Read our brand agreement study for how far the main assistants disagree on the same questions.
If you want a dashboard built this way, see our generative engine optimization service.
What does the research not tell us yet?
The case against raw pooling is strong; the right weights are still an open question. Gaps:
- No study has measured what share of real buyers in a given category use each engine, which is what sensible weights need. AI companies do not publish it.
- The clearest evidence on citation volume covers only Gemini, ChatGPT search and Perplexity, on consumer topics.
- The page-overlap figures come from 15 prompts on one topic in one week, from a single author affiliated with a private firm.
- The opposite-direction example is one brand tracked by a vendor. How often engines diverge for the same brand has not been measured across many brands.
- No study yet compares how well different blended scores predict business results such as traffic or sales.
Frequently asked questions
Should AI visibility be reported per engine or as one number?
Per engine first. Engines cite different volumes and mostly different pages, so a single number is only useful as a labeled summary on top of the per-engine results.
Why does one AI engine dominate my combined visibility score?
Probably because raw counts were added up. Gemini produced up to 50.1 citations per answer in one study, against at most 6.3 for ChatGPT search, so adding counts mostly measures Gemini.
How should we weight ChatGPT, Gemini and Perplexity in one score?
Use a stated rule tied to where your buyers ask, apply it consistently, and show how the score changes under other reasonable weights.
Do ChatGPT and Gemini cite the same sources?
Rarely. In one audit of product questions, the two apps shared only 5.4% of their cited websites for the same question.
Sources
- Sielinski (2026), From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement (opens in a new tab), arXiv:2607.10341.
- Sielinski (2026), Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement (opens in a new tab), arXiv:2603.08924.
- Tannenbaum (2026), Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores (opens in a new tab), arXiv:2609.22655.
- Kumar (2026), Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines (opens in a new tab), arXiv:2606.20065.
- Kato, Honma and Kato (2026), Generative Marketing Mix Modeling: A Causal Inference Framework Linking GEO and GEM to Business Impact (opens in a new tab), arXiv:2609.11915.
- Martinez (2026), Measuring GEO Visibility: Prompt Corpora Define the Answer Market (opens in a new tab), arXiv:2609.06811.
- Uberti-Bona Marin and colleagues (2026), "If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations (opens in a new tab), arXiv:2609.18729.
- Schulte, Bleeker and Kaufmann (2026), Don’t Measure Once: Measuring Visibility in AI Search (GEO) (opens in a new tab), arXiv:2604.07585.
- Underneath (2026), Do ChatGPT, Gemini, Perplexity and Claude agree on brands?