Guide · AI search

How should we measure share of citations in AI search tools?

Measure it over all the answers an engine gives, not just the ones that cite something, and report each engine separately with a range around the figure. Many AI answers never search the web and cite nothing, so a share calculated only among cited answers overstates visibility. And because answers change from run to run, a share from a small sample can rank two sources the wrong way round.

The short version

  1. In one Swiss study, 57.8% of ChatGPT’s runs did not search the web and so cited nothing; a share counted only among the rest would overstate visibility.
  2. In our 80-question study, none of the 42 answers that ran no web search cited anything.
  3. The questions you choose and how you weight them change the result: reweighting the same published data gave rates of 39.7% or 70.5%.
  4. Share estimates from 200 questions on OpenAI’s search still carried ranges of 3 to 6 percentage points, enough to make many apparent leads meaningless.
  5. Engines cite very different numbers of sources, from about 40 to 43 per answer on Gemini to 6 or 7 on OpenAI’s search, so raw counts cannot simply be added up.

Why does the denominator matter so much?

Because many AI answers cite no sources at all, and leaving them out inflates every share you calculate.

Schulte and colleagues (opens in a new tab) ran German-language prompts in four consumer categories on four engines, from Swiss servers, in early 2026. ChatGPT searched the web only for some questions, leaving 57.8% of its runs with no citations. Martinez’s survey (opens in a new tab) draws the lesson: a dashboard “cannot calculate a ‘share of citations’ only among responses that contain citations and then interpret it as overall visibility.” The 57.8% comes from one preprint with a small Swiss question set, so treat it as an example, not a norm.

Our own data shows the same link. In our hidden searches study, every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything. Claude averaged only 0.76 searches per answer, against 3.7 for ChatGPT.

The fix is simple arithmetic. Your share of all answers equals the share of answers that search, times your share of citations within those. Report both parts. A high share among cited answers can coexist with low visibility overall.

Which questions should be in the denominator?

The ones your buyers actually ask, weighted by how often they ask them, and stated openly in the report.

A second paper by Martinez (opens in a new tab) argues that a prompt list defines the market being measured. Stating that a source “is cited in 40% of answers is therefore insufficient” without saying which questions, which engine and which weights. To show the size of the effect, Martinez reweighted three published groups of Google queries. One weighting gave a rate of AI Overviews appearing of 39.7%, another 70.5%, with no change inside any group.

The same applies to citations. A tracker heavy on questions where you are strong will show a high share. Ask any vendor how its question list was built and whether it reflects real buyer demand.

How many answers do we need before a share means anything?

More than most dashboards use, and the number differs by engine.

Sielinski (opens in a new tab), from the measurement firm IQRush, ran 200 questions per topic for nine days on three engines. For most frequently cited sites on OpenAI’s search, the range around each share spanned 3 to 6 percentage points. In one example, one review site showed 9.5% and another 6.0%, yet the ranges overlapped so much that neither could be called the leader. The same problem is why a one-off visibility report can mistake noise for a lead.

Reaching a range of about five points took roughly 40 to 50 questions on Gemini, about 100 on Perplexity and 150 or more on OpenAI’s search. A rise from 8% to 11% on OpenAI’s search, he notes, cannot be credited to your work with confidence.

Engine (Sielinski, 2026)Citations per answerQuestions needed for a range of about 5 points
Geminiabout 40 to 43about 40 to 50
Perplexityabout 20 to 22about 100
OpenAI searchabout 6 to 7150 or more

A follow-up by Sielinski (opens in a new tab) found no fixed budget works everywhere. Across 30 engine-and-topic combinations, no budget below 94 answers would have settled every ranking, and three never settled within 125. Brand counts behave the same way. In our consistency study, pinning a frequency near 50% to within 10 points would take about 97 independent runs.

Can we combine engines into one share of citations?

Only carefully: add raw counts and the engine that cites the most will drown out the others.

Engines cite very different numbers of sources. In Tannenbaum’s four-engine test (opens in a new tab), ChatGPT cited 17.8 pages per prompt and Copilot 4.5. As he puts it, a page competing for a four-source answer faces different odds from one in an 18-source answer. Sielinski warns that pooling raw counts means “a dense platform overshadows sparse ones.”

The better method is to compute a share within each engine, with its own range, and then combine engines using weights you state, such as each engine’s estimated share of your buyers.

Is a citation the same thing as visibility?

No: a citation does not show that anyone noticed it, or that it was accurate.

Citations can misrepresent their sources. In a 2023 audit summarized by Martinez, only 51.5% of sentences in four AI search engines’ answers were fully supported by their citations, and 74.5% of citations supported the claim they were attached to. Kato and colleagues (opens in a new tab) add that how often a name appears is not yet exposure: it must be combined with how many people ask and the chance they notice it.

So report several measures side by side: whether the engine searched, whether you were cited, whether your brand was named, and what was said.

What should you do about it?

Ask for a share that is complete, ranged and broken down by engine, before trusting or buying any figure.

  1. Count every answer, including those with no search and no citations, and show the search rate alongside the share.
  2. Report each engine separately, with its own question count and range.
  3. Use a question list built from real buyer questions, and write down how it is weighted.
  4. Run enough answers per engine: tens on some engines, well over a hundred on others.
  5. Treat changes smaller than the range as noise, and fix the question set before comparing months.
  6. Ask vendors these questions directly: what is the denominator, how many runs, which engine version, which country?

If you want help building a measurement that meets these tests, see our generative engine optimization service.

What does the research not tell us yet?

It does not yet give a standard sample size or tell us how citation share relates to sales.

  • The 57.8% no-search figure comes from one small Swiss preprint; rates elsewhere are unknown.
  • Sample-size guidance comes from one vendor’s data across three engines; its own author says thresholds may not transfer.
  • Engines change their search behavior often, so today’s search rates may not hold.
  • No study here links citation share to clicks, leads or revenue.
  • How to weight engines by their real audience has not been studied.

Frequently asked questions

What is share of citations in AI search?

It is the fraction of the sources AI engines cite that point to your site. It means little unless you state the questions, the engine and whether answers with no citations were counted.

Why do AI visibility tools show different numbers?

Because they use different question lists, engines, run counts and denominators. Reweighting the same published data alone moved one rate from 39.7% to 70.5%.

How many prompts do we need to track AI citations reliably?

It depends on the engine. One study needed about 40 to 50 questions on Gemini but 150 or more on OpenAI’s search for a range of about five points.

Should answers without citations count in our share?

Yes. In one study, 57.8% of ChatGPT runs searched nothing and cited nothing; dropping them overstates visibility.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.