The short version
- Brand mentions are steadier than cited pages. In a Swiss study, brand lists overlapped 45% to 59% from one day to the next, cited sources only 34% to 42% (Schulte and colleagues, 2026).
- Being cited is not the same as being used. ChatGPT cited 6.88 sources per question against Perplexity’s 16.35, yet each page it cited shaped its answer far more (Zhang Kai and colleagues, 2026).
- Being known is not being recommended. In one study ChatGPT recognized 99.4% of products when asked by name but surfaced only 3.32% in open discovery questions (reported by Martinez, 2026).
- Each engine is its own market: in one small audit, 96.4% of cited web addresses appeared on only one of four engines (Tannenbaum, 2026).
- Presence is not accuracy: in our pricing study, 61.9% of software plan prices quoted by AI assistants were fully faithful to the vendor’s page.
Why measure final answers instead of rankings or page scores?
Because the final answer is what people read. Rankings and page scores only describe steps before it.
An AI answer is built in stages. The engine decides whether to search, runs its own searches, picks some pages, writes an answer and attaches citations. A page can drop out at any stage. Wen and colleagues (opens in a new tab) argue that academic studies often report movement in intermediate ranked lists. By contrast, how often a source appears and is cited in the final answer shows whether it survived every stage.
Google rankings are a weak stand-in. In our rankings study, only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the question asked. Tannenbaum (opens in a new tab) makes the same point about scores that rate a page without ever querying an engine. In his audit of 15 prompts on one day, 96.4% of the web addresses cited appeared on only one of ChatGPT, Copilot, Google and Perplexity. A page score cannot tell you which engine will find you. The author works for a company that sells AI visibility scoring.
Should you track brand mentions or cited pages?
Track both, but treat brand mentions as the headline. They are steadier, and most citations point at someone else’s site.
Schulte and colleagues (opens in a new tab) followed four AI engines in four Swiss industries for about 45 days. From one day to the next, the brands named overlapped 45% to 59%, while the sources cited overlapped only 34% to 42%. They conclude that brand presence, added up over a campaign, is a more reliable measure than individual cited web addresses. Their brand detection used a fixed list of names, and one industry was dropped because answers rarely named brands.
Cited pages and named brands can also move independently. In our consistency study, Perplexity cited identical web addresses in 166 pairs of runs, yet the brand list still changed in 91.6% of them. And your own site is a small slice of the evidence. On one tracking platform, only 2.9% of citations pointed at the tracked brand’s own domain, while 75.2% pointed at other companies in the same space. That figure comes from Kumar (opens in a new tab), a co-founder of the platform.
Is being cited the same as being used?
No. A page can be listed as a source while contributing little to what the answer actually says.
Zhang Kai and colleagues (opens in a new tab) studied 602 prompts across ChatGPT, Google and Perplexity. They scored how much each cited page shaped the answer, from repeated references to shared wording.
| AI engine | Citations per question | How much each cited page shaped the answer (0 to 1) |
|---|---|---|
| ChatGPT | 6.88 | 0.2713 |
| Google AI Overview or Gemini | 12.06 | 0.0584 |
| Perplexity | 16.35 | 0.0646 |
The authors’ advice is to track both whether you are cited and how much your page shapes the answer. A dashboard that counts citations alone misses the difference. Their influence score is their own construction from one static dataset, not a view inside the engines.
Citations also do not guarantee support. In a 2023 audit of four AI search engines, summarized by Martinez (opens in a new tab), only 51.5% of sentences in AI search answers were fully supported by their citations. Only 74.5% of citations supported the statement they were attached to.
Does it matter whether the question names your brand?
Yes, enormously. Questions that name you measure recognition; open questions measure whether you get recommended.
Martinez reports a preprint by Sharma on 112 startups. ChatGPT recognized 99.4% of products when they were named, but surfaced them in only 3.32% of open discovery questions. For Perplexity the drop was from 94.3% to 8.29%. That is one preprint using two models, but the distinction it draws matters for any score. A visibility score padded with branded questions will look healthy while customers never hear your name.
The denominator matters too. In the Swiss data, 57.8% of ChatGPT runs did not search the web at all. A share of citations computed only among answers that cite something overstates how often people see you.
Do position and tone matter as well as presence?
Yes. Two answers can both name you while one puts you first and the other sixth.
In our consistency study, 86.4% of the brands ChatGPT named in every one of five runs still moved position at least once. Its first pick changed at least once for 80.0% of questions. So track separately how often you are named, how often you are in the top three, and how often you are first. Engines differ here too; see which AI engine is most consistent.
Tone is harder to measure reliably. On the platform studied by Kumar, the tone toward a brand flipped between positive and negative in 45.5% of cases. Whether the brand was mentioned flipped in only 6.8%. Tone readings need many more answers before they settle, and the platform’s own model scored the tone.
Should you check whether the answer is accurate?
Yes. Being named with the wrong price, phone number or claim can cost you more than not being named.
In our pricing study, 61.9% of plan prices quoted by four AI assistants were fully faithful to the vendor’s page. In our business facts study, 18.9% of answers about local businesses stated at least one fact that differed from the business’s Google profile. Not every difference was an error, which is why the check needs a human eye.
What should you do about it?
Build a small set of measures that follow the answer from search to sale. In practice:
- Lead with brand mention rate on open questions that do not name you, per engine, over repeated runs.
- Report branded questions separately. They show recognition, not recommendation.
- Track position: share of answers where you are first, in the top three, or named at all.
- Track citations of your own pages as a second measure, and note which third-party pages cite you.
- Count answers with no search or no citations in the denominator rather than dropping them.
- Audit accuracy of prices, facts and claims in a sample of answers every month.
- Keep engines separate. A gain on one engine does not carry over to another.
If you want a measurement plan built this way, see our generative engine optimization service.
What does the research not tell us yet?
No study yet links any of these measures to sales or revenue. Other gaps:
- Which measure predicts business results. Wen and colleagues argue final-answer measures sit closer to customer attention, but they did not measure clicks or revenue.
- Whether influence scores reflect real use. Measures of how much a page shaped an answer are built from text overlap, not from the engines’ internals.
- Whether customers notice citations. Few studies observe whether people see or click the sources listed.
- How stable these patterns are. Most findings come from one period, a few markets and a handful of engines, and the engines change often.
Frequently asked questions
What is the best single metric for AI visibility?
There is no single best one, but brand mention rate on unbranded questions, per engine, is the most stable starting point. In the Swiss study, brand lists overlapped 45% to 59% from day to day, against 34% to 42% for cited sources.
Is share of citations a good measure of AI visibility?
It is useful but incomplete. Citations count exposure, not influence, and ChatGPT cited 6.88 sources per question while using each far more heavily than Perplexity used its 16.35.
Should I track prompts that name our brand?
Yes, but report them separately. Engines recognize named brands almost every time, so branded prompts inflate a visibility score that should reflect open recommendation questions.
Do Google rankings tell me how visible I am in AI answers?
Not reliably. Only 8.3% of the pages ChatGPT cited ranked in Google’s top 10 for the same question in our study.
Sources
- Wen, Y., Zhang, N., Yuan, H., Chen, X., Zhang, H. and Guo, H. (2026), Position: Generative Engine Optimization Creates Underexamined Risks, Governance Must Target Concentration, Disclosure, and Academic Blind Spots (opens in a new tab), arXiv:2606.12439.
- Schulte, J., Bleeker, M. and Kaufmann, P. (2026), Don’t Measure Once: Measuring Visibility in AI Search (GEO) (opens in a new tab), arXiv:2604.07585.
- Zhang Kai, He Xinyue and Yao Jingang (2026), From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms (opens in a new tab), arXiv:2604.25707.
- Martinez, O. (2026), Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026) (opens in a new tab), arXiv:2607.14035.
- Tannenbaum, B. (2026), Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores (opens in a new tab), arXiv:2609.22655.
- Kumar, P. (2026), Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines (opens in a new tab), arXiv:2606.20065.
- Underneath (2026), Ask an AI the same question 5 times: do the brands change?
- Underneath (2026), Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?
- Underneath (2026), How faithfully do AI assistants quote software prices?
- Underneath (2026), Do AI answers match a business’s Google profile?