The short version
- In a test of 20 brands, a single AI answer could barely rank brands at all (about 0.01 on a 0 to 1 reliability scale), and even 8 languages, 3 models and 15 phrasings reached only about 0.36 (Żatuchin, 2026 (opens in a new tab)).
- In the same study, 15 phrasings asked once each (360 queries) ranked brands better than 5 phrasings asked 5 times each (600 queries).
- A Swiss study of four AI search engines recommends at least 7 runs per prompt per day for brand tracking, pooled over two to four weeks (Schulte and colleagues, 2026 (opens in a new tab)).
- Weighting the same published data two ways gave AI answer rates of 39.7% and 70.5%, which shows how much the choice of questions drives a headline score (Martinez, 2026 (opens in a new tab)).
What are you really measuring when you track AI visibility?
You are measuring your brand against a market you defined yourself: your list of questions and how much each counts. Martinez (opens in a new tab) calls this the “answer market”. A score of “cited in 40% of answers” means little until you know which questions, which assistant, which period and which weights produced it.
The paper shows how much this matters with published data on Google’s AI Overviews. It took activation rates for three groups of questions and weighted them two different ways, without changing any rate inside a group. The result was an AI answer rate of 39.7% under one weighting and 70.5% under the other. That example concerns whether an AI answer appears, not brand visibility, but the lesson carries over.
Two practical points follow. A question set that leans heavily on easy, branded questions will flatter you. And adding many rewordings of one question quietly gives that question more weight unless you correct for it.
Where does the noise in AI visibility scores come from?
At least four places: chance, wording, which assistant you ask and which language you ask in. Żatuchin (opens in a new tab) measured all four on 12,933 answers about 20 Central and Eastern European brands, in eight languages, from three AI models. The point of the design was to see which source of noise a tracking budget should target.
Most tracking tools spend their budget on repeats of the same prompt, typically five. The paper argues that repeats tackle only one source of noise, and the cheapest one: averaging over every query already shrinks it. Language and assistant differences shrink only when you add more languages and assistants.
The author is affiliated with Rankfor.AI, which sells brand-visibility measurement, and says the method is meant for sizing its own measurements. The study ran in one window in spring 2026 and measured the tone of answers rather than whether brands were named.
Is it better to add more runs or more variety?
Variety wins, by a wide margin in this test. Starting from a common design of one language, one model, five phrasings and five repeats, the study compared where the next queries should go. Three more languages cut measurement error about fifteen times as much as five more repeats. The language option used three times as many extra queries, yet it still came first per query. More models and more phrasings came next; repeats came last.
Piling on repeats barely helped. Twenty repeats of one prompt in one language lifted the brand-ranking reliability score only to 0.020. A design of 15 phrasings asked once each, 360 queries in total, scored higher than one of 5 phrasings asked 5 times each, at 600 queries. Breadth beat depth even while costing less.
The ceiling was low everywhere. On a 0 to 1 scale, a single answer scored about 0.01, and the full design of 8 languages, 3 models and 15 phrasings reached about 0.36. The models were run at a low randomness setting, so real chat apps may vary more between repeats than this study saw. That is why a single AI visibility snapshot is a weak basis for decisions.
How many runs and questions do you need?
There is no universal number; it depends on the assistant and the topic. Still, several studies give useful floors.
| Study | Setting | Suggested minimum |
|---|---|---|
| Schulte and colleagues (opens in a new tab) | 4 engines, Swiss-German prompts, early 2026 | At least 7 runs per prompt per day for brands, 8 for sources, pooled over two to four weeks |
| Sielinski (opens in a new tab) | 3 engines, 3 consumer topics, 9 days | About 40 to 50 questions on Gemini, about 100 on Perplexity, 150 or more on ChatGPT search, to pin citation shares to a range about five points wide |
| Sielinski (opens in a new tab) | 3 engines, 10 topics, 125 questions each | No budget below 94 responses was enough for every topic; three topics were not settled within 125 |
| Our consistency study | 3 assistants, 20 questions, 5 runs | About 97 runs to pin a brand’s naming rate near 50% to within 10 points |
Sielinski is affiliated with the company IQRush and Schulte with Aurora Intelligence, so treat these as practitioner studies. The first Sielinski paper warns against stopping as soon as the numbers look steady, because the error can narrow and then widen again. Its advice is to fix the number of questions in advance, based on earlier measurements of that assistant and topic.
Which wordings belong in your prompt set?
Wordings that real buyers use, chosen on purpose, because a different wording is often a different question. In our rewording study, asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it only 15.3% of the time. That change is not noise; the buyer asked for something else.
So separate two kinds of variation. Rewordings that keep the need the same, such as “best” versus “top-rated”, measure how robust your visibility is. Rewordings that add a need, such as a budget, a company size or a region, are new questions that deserve their own line in the report.
Individual prompts also behave very differently. In the Swiss study, some prompts returned almost the same sources run after run, with overlap above 0.8, while others stayed below 0.2. A score built on one or two prompts mostly reflects the quirks of those prompts.
Should you track the consumer app or the developer API?
Track what buyers see, which usually means the consumer app, or label clearly that you measured the API. In a Dutch audit of product questions (opens in a new tab), ChatGPT’s app and its API, asked the same question moments apart, shared only 12.0% of their cited domains on average. Our guide on monitoring visibility through the API covers that gap in detail.
Assistants also differ in how often they search at all. In the Swiss study, ChatGPT left 57.8% of its runs with no citations because it only searched for some questions. If you track citations, budget for those empty runs. Sielinski (opens in a new tab) gives a worked example: needing 100 answers with citations when 83% of queries return them means submitting about 121 queries. Our guide on measuring share of citations fairly explains how to count those empty runs.
What should you do about it?
Design the tracking program before you buy the tool, and write the design down. Steps:
- Define the answer market: list the buyer questions, the markets and the weight each should carry, and keep branded and unbranded questions separate.
- Cover the languages and locations your buyers use, then the assistants they use, before adding repeats.
- Use several real-buyer wordings per need, and report need-changing wordings such as budget or company size as separate questions.
- Fix the run count in advance. As a floor, use the published minimums above, and pool results over two to four weeks rather than reading daily swings.
- Measure the consumer app where you can, and record the date, location, assistant and settings with every run.
- Report each result as a rate with its range, such as “named in 6 of 10 runs”, per assistant and per market.
For help designing a tracking program like this, see our generative engine optimization service.
What does the research not tell us yet?
The research agrees on the direction but cannot yet give you an exact recipe. Open questions:
- The clearest budget comparison comes from one study of 20 brands in one region and one window, measuring tone rather than whether brands were named. Its author plans a version that measures naming.
- Most studies here were written by people tied to AI-visibility vendors. Their methods are transparent, but independent replication is limited.
- Suggested run counts come from specific topics and markets: Swiss-German consumer queries, US consumer products, Central European brands. B2B, local and regulated categories are largely untested.
- No study has yet tied a prompt set to real buyer question volumes, because AI companies do not publish them. Every prompt set is still an informed guess about demand.
Frequently asked questions
How many times should I run each prompt to track AI visibility?
One Swiss study recommends at least 7 runs per prompt per day for brand tracking, pooled over two to four weeks. Another found repeats past five add little once you cover more languages, assistants and wordings.
How many prompts do I need for AI visibility tracking?
It depends on the assistant. One study needed about 40 to 50 questions on Gemini, about 100 on Perplexity and 150 or more on ChatGPT search before citation shares settled to a range about five points wide.
Should AI visibility prompts include different phrasings?
Yes. In one test, 15 phrasings asked once each ranked brands better than 5 phrasings asked 5 times each, while using fewer queries.
Can I track AI visibility through the API instead of the app?
You can, but label it as API data. In one audit, ChatGPT’s app and API shared only 12.0% of their cited domains for the same question asked moments apart.
Sources
- Żatuchin (2026), Where Does the Noise Come From? A Variance-Components Decomposition of Non-Determinism in LLM Brand Answers (opens in a new tab), arXiv:2607.13304.
- Schulte, Bleeker and Kaufmann (2026), Don’t Measure Once: Measuring Visibility in AI Search (GEO) (opens in a new tab), arXiv:2604.07585.
- Sielinski (2026), Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement (opens in a new tab), arXiv:2603.08924.
- Sielinski (2026), From Stochastic to Stable: Rank Stability and Structural Sufficiency in AI Visibility Measurement (opens in a new tab), arXiv:2607.10341.
- Martinez (2026), Measuring GEO Visibility: Prompt Corpora Define the Answer Market (opens in a new tab), arXiv:2609.06811.
- Uberti-Bona Marin and colleagues (2026), "If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations (opens in a new tab), arXiv:2609.18729.
- Underneath (2026), Ask an AI the same question 5 times: do the brands change?
- Underneath (2026), Does rewording a question change AI brand recommendations?