Guide · AI search

How many prompts do we need to track to measure AI visibility?

There is no single right number, and anyone who quotes one is guessing. The best current research points to roughly 40 to 150 or more prompts per topic for each AI engine, each asked several times. The exact figure depends on the engine, the topic and how fine a difference you need to see. Which prompts you choose matters as much as how many.

The short version

  1. Across three consumer topics, Gemini needed about 40 to 50 prompts per topic for a tight reading, Perplexity about 100 and OpenAI’s search model 150 or more (Sielinski, 2026).
  2. In a second dataset of 30 engine-and-topic pairs, no budget below 94 answers would have been enough for every pair, and three pairs were still unsettled at 125 (Sielinski, 2026).
  3. A Swiss study across four AI engines recommends asking each prompt at least 7 times a day to estimate whether a brand appears (Schulte and colleagues, 2026).
  4. In our own test, a single ChatGPT answer showed only 57.8% of the brands that five answers to the same question named between them (our consistency study).
  5. The mix of prompts can swing a score on its own: reweighting the same published data produced AI answer rates of 39.7% or 70.5% (Martinez, 2026).

Is there a standard number of prompts?

No. The research shows the right number changes with the engine and the topic, so no fixed count holds everywhere.

The clearest test comes from Sielinski (opens in a new tab), who tracked 10 topics on Gemini, Perplexity and OpenAI’s search model. He asked when each ranking of cited websites stopped moving and became precise enough to act on. The method settled for 27 of the 30 engine-and-topic pairs within 125 answers. In that data, no budget below 94 answers would have covered every pair that settled, and three pairs had still not settled by 125.

Early answers are especially misleading. In the same paper, the ranking of the most-cited websites kept changing substantially through the first 20 to 60 queries, depending on topic and engine. The author works for a company that sells AI visibility measurement, and the queries were written by an AI tool rather than taken from real users.

How many prompts does each AI engine need?

Fewer for Gemini, more for Perplexity, and the most for ChatGPT’s search model. The gap tracks how many sources each engine cites per answer, and how erratically.

In an earlier paper, Sielinski (opens in a new tab) sent the same 200 queries per topic to the three engines every day for nine days, across bird feeders, multivitamins and running gear. He then asked how many queries it takes to narrow a website’s share of citations to a range about five percentage points wide.

AI enginePrompts per topic for a five-point reading
GeminiAbout 40 to 50
PerplexityAbout 100
OpenAI search model (called SearchGPT in the study)150 or more, and sometimes no fixed number worked

There is a hidden cost on top. OpenAI’s model sometimes answers without citing anything: on one topic in the later paper, it returned citations for only 104 of 125 questions, a zero-citation rate of about 17%. To end up with 100 usable answers, the author says you would need to send about 121 questions rather than 100. Our guide on measuring share of citations explains how to count those empty answers.

How many times should each prompt be asked?

More than once, and probably at least seven times. The same prompt gives a different answer from one run to the next.

Schulte and colleagues (opens in a new tab) ran eight prompts per industry through ChatGPT, Gemini, Google AI Mode and Perplexity in four Swiss industries over 45 days. From repeated same-day runs, they concluded that brand tracking needs at least 7 runs per prompt per day, and at least 8 if you care which web pages are cited. One author is affiliated with the company that supplied the data, and the thresholds assume brands that appear some of the time, not always or never.

Our own test points the same way. We asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each. One ChatGPT answer showed 57.8% of the brands that the five answers named between them. Even five runs is coarse: pinning a brand that appears about half the time to within 10 points would take about 97 runs (our consistency study). That is why a one-time AI visibility report can mislead.

Does the mix of prompts matter more than the count?

Often, yes. Which questions you ask, in what words, language and country, can move a score more than adding runs.

Martinez (opens in a new tab) argues that a prompt set defines its own “answer market”, which need not match what real customers ask. Saying a source is cited in 40% of answers means little, he notes, until you know what was asked. Open questions, questions naming the brand and comparisons give very different rates. Using published figures for Google, he showed that two weightings of the same three groups of searches produced AI answer rates of 39.7% and 70.5%, with nothing else changed. That calculation concerns whether an AI answer appears, not brand visibility, but the lesson carries over.

Is it better to add new prompts or repeat old ones?

Adding different prompts usually buys more precision than repeating the same ones. Repeats reduce only one kind of noise.

A study of 12,933 answers about 20 Central and Eastern European brands split the variation in the tone of AI answers into its sources (Żatuchin (opens in a new tab)). A design that spread 360 questions across 15 wordings measured brands more reliably than a 600-question design that asked 5 wordings 5 times each. The outcome was the tone of each answer, not whether a brand was named, and the author sells brand measurement, so treat the exact sizes with care.

Schulte and colleagues reached a similar view by another route. Within one industry, some prompts gave nearly the same sources every run, with overlap scores above 0.8, while others stayed below 0.2. Tracking one or two prompts would mostly measure the quirks of those prompts.

How long should you track before trusting a trend?

Weeks, not days. Daily snapshots of a single brand stay noisy for several weeks of collection.

In the Swiss data, the estimate for one brand’s appearance rate only became reasonably steady after about 24 days of daily tracking. The authors recommend rolling averages over two to four weeks. Sielinski found a related trap: rankings can look steady from one day to the next while drifting over nine days. The same caution applies when you judge whether a GEO campaign worked.

What should you do about it?

Start with the decision the numbers must support, then size the prompt set to it. A practical plan:

  1. Set the precision you need first. Telling a brand at 12% from one at 10% needs far more prompts than spotting that a competitor appears twice as often as you.
  2. Size per engine. Plan for about 40 to 50 prompts per topic on Gemini and 100 or more on Perplexity and ChatGPT’s search, then check whether your rankings have settled.
  3. Send extra prompts where answers lack citations. If about one answer in six comes back with no sources, as in one test, send about a fifth more prompts.
  4. Repeat each prompt several times. Seven runs is a reasonable starting point for brand mentions.
  5. Build the set to match real demand. Mix broad and specific questions, budget and premium wording, and every country and language you sell in. Report each group separately.
  6. Report a range, not a single figure. “Named in 4 of 7 runs” is honest; “visible in ChatGPT” is not.
  7. Judge trends on two-to-four-week averages, not day-to-day moves.

If you want help designing a prompt set like this, see our generative engine optimization service.

What does the research not tell us yet?

The research has not produced a tested, general rule for how many prompts a brand needs. Specific gaps:

  • Real user questions. The main studies used prompts written by an AI tool or taken from search suggestions, not logs of what customers actually ask.
  • Business-to-business topics. The sample-size work covered consumer topics such as running gear and smoke detectors.
  • Brand mentions on every engine. The per-engine figures above measure how often websites are cited, not how often a brand is named.
  • Long-term drift. Most datasets cover days or weeks, so model updates over months are not captured.
  • Independent replication. Several key studies come from companies that sell measurement, and none has been repeated by an outside team.

Frequently asked questions

Is 50 prompts enough to measure AI visibility?

It can be for one topic on Gemini, but probably not on ChatGPT’s search or Perplexity. In Sielinski’s test, those engines needed about 100 and 150 or more prompts per topic for a similar precision.

Should I track the same prompts every day?

Yes, keep a fixed core set so you can compare over time, and repeat each prompt several times. Schulte and colleagues recommend at least 7 runs per prompt per day for brand tracking.

Why does my AI visibility score jump around from week to week?

Mostly because each answer is a sample from a changing distribution. A single ChatGPT answer showed only 57.8% of the brands five answers named in our test, so small prompt sets swing a lot.

Do I need prompts in other languages?

Yes, if you sell in other languages. For local European brands, asking in the home language raised how often they were recommended by 0.80 on a 0-to-1 scale.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.