Guide · AI search

Can I trust a one-time AI visibility report for my brand?

Not on its own. AI assistants give a different answer each time, so a report built from one round of questions can rank you ahead of a competitor, or behind, purely by chance. Researchers who have measured the noise recommend repeated runs, several wordings and several engines, tracked over two to four weeks, before a number is solid enough to drive a budget decision.

The short version

  1. In our test of three assistants, a single ChatGPT answer showed only 57.8% of the brands that five answers to the same question named between them.
  2. A visibility lead of 9.5% against 6.0% turned out to be statistically indistinguishable from a tie once the measurement’s margin of error was added (Sielinski, who works for an AI visibility company).
  3. Swiss researchers recommend at least 7 runs per question per day and rolling results over two to four weeks before trusting a brand-level number.
  4. Repeating one question helps less than asking it in more languages, on more engines and in more wordings, according to a study of 20 brands in 8 languages.
  5. One vendor’s data suggest a single check is right more often for clear cases: 63.2% of brand-and-question pairs were never mentioned in any run.

Why can’t one check tell you where your brand stands?

Because the same question asked twice gets a different answer, so one answer is only a sample. In our consistency study, we asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. Of the brands ChatGPT named for a question across five runs, 25.2% appeared in all five and 36.6% in only one.

A single ChatGPT answer showed 57.8% of the brands its five answers named between them; for Gemini, 48.4%. ChatGPT’s first pick changed at least once for 80.0% of questions. How steady each engine is on its own is compared in which AI engine answers most consistently. Timing mattered too: an answer to the same question about 4.4 hours later overlapped less with the earlier runs than they did with each other.

Schulte and colleagues (opens in a new tab) put it plainly after tracking four engines for 45 days. A snapshot “may differ substantially from a second query executed minutes later under identical conditions.” Malthouse and colleagues (opens in a new tab) at Northwestern University reach the same view: a recommendation should be treated as a set of probabilities, not a single list.

How big is the noise compared with the differences you care about?

Often bigger than the gaps reports highlight. Sielinski (opens in a new tab) gives a typical case from ChatGPT search. In a sample of 200 running-gear questions, tomsguide.com held a 9.5% share of citations and runnersworld.com 6.0%. The first looks like the clear leader.

Add the margin of error and the picture changes. The plausible range ran from 5.5% to 12.5% for the first site and from 4.0% to 8.0% for the second, so the ranges overlap and the lead may be noise. On ChatGPT search, those ranges were typically 3 to 6 points wide. By the same logic, a rise from 8% to 11% after a content change cannot be credited to the change. Telling a real gain from noise is covered in judging whether a GEO campaign worked.

Single samples can also mislead badly. One site took a share of 0.032 of Gemini’s citations in the first daily sample, against a long-run average of 0.005 across the following days. Anyone reading only the first day would have ranked it among the top sources.

How many checks does a reliable number need?

More than most one-off reports use, and the right number depends on the engine and topic. Schulte’s team found brand-level figures needed at least 7 runs per question per day, and source-level figures at least 8. Over time, a brand’s figure became reasonably steady after about 10 days and tight after about 24 days.

Sielinski estimated how many questions it takes to pin a domain’s citation share within a five-point range. Gemini needed roughly 40 to 50 questions, Perplexity about 100 and ChatGPT search 150 or more.

What you want to knowWhat the research suggests
Whether a brand appears for one questionAt least 7 runs per day (Swiss study)
A share accurate to about 10 pointsAbout 97 runs (our consistency study)
A citation share accurate to 5 points40 to 150 or more questions, by engine (Sielinski)
A stable trend for one brandTwo to four weeks of rolling results (Swiss study)

Five runs leave wide margins. In our study, a brand named in 3 of 5 runs could plausibly appear anywhere from 23.1% to 88.2% of the time. In a later paper, Sielinski found that the data needed before rankings settled, under one common test, ranged from 40 responses to never across 30 engine and topic combinations.

Is asking the same question more times enough?

No: the questions you choose and the engines you ask matter as much as repeat runs. Żatuchin (opens in a new tab), who is affiliated with the AI visibility company Rankfor.AI, analyzed answers about 20 Central and Eastern European brands in 8 languages on 3 engines. The language of the question explained 26.5% of the variation in a single answer; the brand itself explained 1.5%.

On a 0-to-1 scale of how reliably brands could be ranked, one answer scored near 0.01. Even 8 languages, 3 engines and 15 wordings together reached only about 0.36. The outcome measured was the tone of answers, and 91.9% of answers were neutral, so the figures may not carry over to recommendations.

Wording matters on its own. In our phrasing study, simply asking again kept the same first brand 68.0% of the time; adding “on a tight budget” kept it 15.3%. Martinez (opens in a new tab) shows how much the choice of question set can move a headline figure. Reweighting the same published data gave AI answers appearing for 39.7% or 70.5% of searches, depending only on how the question groups were weighted.

Can a single check ever be trusted?

Sometimes, for clear-cut cases: a brand that is never or always named is easy to spot. Kumar (opens in a new tab), a co-founder of the tracking company Ranqo, analyzed more than 100 brands tracked between March and May 2026. Of brand, question and engine combinations, 77.5% were always or never mentioned across runs, and only 6.8% flipped often.

The most common result was absence: 63.2% of combinations were never mentioned at all. For the rest, Kumar recommends at least 3 runs before trusting a mention rate. A one-off check can therefore tell you that you are missing, but it is a poor guide to how often you appear when you sometimes do.

Why does tracking need to be continuous?

Because the engines change, and a brand’s position can drift without anyone noticing. Kumar’s data show that brands which did nothing slowly lost visibility on ChatGPT and Perplexity, by an average of 1.34% per tracking run on ChatGPT, while holding steady on the other engines. Kumar notes that part of this may be measurement drift.

Baig and colleagues (opens in a new tab), who audited how AI models choose hotels, end with a warning for managers. The weights they measured describe the model versions of the day, and “the appropriate managerial posture is continuous measurement rather than a one-time fix.”

What should you do about it?

Treat a one-off AI visibility report as a starting hypothesis, and ask how it was measured before acting on it.

  1. Ask how many times each question was run, on how many days, and on which engines. One run per question is a sample of one.
  2. Ask for margins of error. A lead that sits inside the margin is a tie.
  3. Use a broad, fixed set of questions in several wordings, and in each language your buyers use.
  4. Track over at least two to four weeks before comparing yourself with competitors or judging a change.
  5. Act quickly on clear absences. If you are never named across repeated runs, that finding is reliable.
  6. Keep measuring after any change, because engines update and early gains can fade.

For help building tracking that holds up, see our generative engine optimization service.

What does the research not tell us yet?

There is no agreed standard for how much AI visibility data is enough.

  • The recommended run counts come from small settings: German questions on Swiss servers, three consumer topics, or 20 regional brands.
  • Several of the key studies come from authors who sell tracking tools, which gives them a stake in the case for repeated measurement.
  • Żatuchin measured the tone of answers, not recommendations; a version using recommendations is still to come.
  • No study has yet shown how much a measurement error costs a brand in lost sales.

Frequently asked questions

How many times should I run a prompt to measure AI visibility?

At least 7 times per day per question, according to a 45-day Swiss study. Precise shares take far more: pinning one to within about 10 points took about 97 runs in our estimate.

How long should I track AI visibility before judging results?

Two to four weeks of rolling results, according to the Swiss study. A brand’s figure became reasonably steady after about 10 days and tight after about 24.

Is an AI visibility score from a single audit useful?

Yes for spotting clear absences, much less for rankings. In one vendor dataset, 63.2% of brand and question combinations were never mentioned in any run, while close rankings often fell within the margin of error.

Why did my brand’s AI visibility change between two reports?

Possibly only chance. In one example, a site at 9.5% and another at 6.0% had overlapping margins of error, so a change of a few points can be noise.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.