---
title: "Ask an AI the same question 5 times: do the brands change?"
description: "Same 20 buyer questions, five runs each: a quarter of ChatGPT’s brands appeared every time. AI visibility is a distribution, and one answer is a sample."
canonical: "https://underneath.agency/research/ai-recommendation-consistency-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Ask an AI the same question 5 times: do the brands change?

When a report says a brand is “visible in ChatGPT” for a question, it usually means the brand appeared in one answer. But assistants give a different answer each time. We asked ChatGPT, Gemini and Perplexity the same 20 buyer questions five times each on 26 September 2026, with the five runs of a question a median of about 11 minutes apart, and compared all 300 answers.

Version 1.0 of this study showed that the brands change. This version asks what that means for measurement. We treat each brand’s visibility as a frequency: the share of runs that name it. Then we look at how the change is structured (which brands, in what order, which one first, from which sources), what it goes with, and how much a single answer gets wrong. No new questions were asked; the 300 answers are the same, plus the answers to the same questions from our country and rewording studies, collected a few hours later that day.

The short answer: AI visibility is not a yes-or-no property of a brand and a question. It is a distribution with a small core and a long tail, and its shape depends on the assistant, the question and when you ask.

## The short version

1. A quarter of ChatGPT’s brands were named in every run. Of the brands ChatGPT named for a question across five runs, 25.2% appeared in all five (95% interval 15.9% to 36.8%) and 36.6% in only one. For Gemini the figures were 13.7% and 46.7%; for Perplexity 40.9% and 15.8%.
2. One answer shows about half the picture. A single ChatGPT answer showed 57.8% of the brands its five answers named between them; Gemini 48.4%, Perplexity 72.2%. The fifth run still added about one new brand on ChatGPT and Gemini, so five runs had not found the whole tail either.
3. Being named is not the same as being ranked. ChatGPT named the same first brand in two runs of a question 53.8% of the time, and its first pick changed at least once for 80.0% of questions. Of the brands ChatGPT named in every run, 86.4% moved position between runs.
4. Same sources, different brands. In 166 pairs of Perplexity runs the cited URLs were identical, yet the brand list changed in 91.6% of them. Across all three assistants, run pairs that shared more sources shared slightly more brands (about 0.03 more brand overlap per 0.1 more source overlap), but source overlap explained little of the change in brands (correlation 0.357 on ChatGPT, 0.228 on Gemini).
5. Stability depends on the question as much as the assistant. The mean overlap between runs ranged from 0.299 to 0.747 across questions. The question accounted for 30.0% of the variation, the assistant for 25.5%, and the combination of the two for the rest; a question that was stable on one assistant was not reliably stable on another.
6. Time matters within hours. For ChatGPT, an answer to the same question about 4.4 hours later overlapped with the five runs less (0.477) than the runs overlapped with each other (0.588), and an answer for another country less again (0.398). Rewording a question changed Gemini’s brands more than asking again did; for ChatGPT, most of the apparent rewording effect was already there when the same question was asked hours later.
7. Five runs are too few for a precise number. A brand named in 3 of 5 runs has a 95% interval of 23.1% to 88.2%. Pinning a frequency near 50% to within 10 points would take about 97 independent runs; our data cannot yet say how many are needed in practice.

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. How much do the recommended brands change across identical runs? | Yes |
| RQ2. How is the change structured: core and tail, order, first pick? | Yes |
| RQ3. Is the change in brands tied to a change in cited sources? | Yes, as an association |
| RQ4. What does one run miss, and how precise are five? | Partly |
| RQ5. How does repetition compare with hours later, rewording and country? | Partly, on 5 to 10 questions |
| RQ6. Drift over days and weeks; runs needed in practice; retrieval vs generation | No, planned |

## A brand’s visibility is a frequency, not a yes or no

Across five runs, each brand an assistant named for a question gets a frequency: named in 1, 2, 3, 4 or 5 of the runs. We group them as core (5 of 5), stable (4 of 5), intermittent (2 or 3) and tail (1 of 5).

| Share of brands | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Core (5 of 5) | 25.2% | 13.7% | 40.9% |
| Stable (4 of 5) | 15.3% | 14.1% | 14.8% |
| Intermittent (2 to 3) | 22.9% | 25.6% | 28.6% |
| Tail (1 of 5) | 36.6% | 46.7% | 15.8% |
| Brand-question pairs | 262 | 227 | 203 |

The 95% intervals, from resampling the 20 questions, are wide for the core share: 15.9% to 36.8% for ChatGPT, 8.3% to 19.8% for Gemini and 35.0% to 47.8% for Perplexity. The ordering of the three assistants holds: Perplexity’s mean run-to-run brand overlap was 0.135 higher than ChatGPT’s (interval 0.034 to 0.247) and 0.260 higher than Gemini’s (0.168 to 0.359), and ChatGPT’s was 0.125 higher than Gemini’s (0.013 to 0.233).

The frequency is an observed share of five runs, not a known probability. With five runs it can only place a brand in broad bands (see “How precise five runs are” below).

## How concentrated the recommendations are

Counting the brands in one answer understates how many brands an assistant spreads its answers over. We measured the effective number of brands per question: the number of equally frequent brands that would give the same spread (the exponential of the entropy of each brand’s share of all mentions across the five runs).

| Per question (median) | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Brands in one answer | 6.5 | 5.5 | 6.8 |
| Distinct brands in 5 answers | 11 | 10.5 | 9.5 |
| Effective number of brands | 9.5 | 9.2 | 8.6 |
| Core brands (5 of 5) | 3 | 2 | 4 |
| Tail brands (1 of 5) | 4 | 4 | 1 |

On average Gemini’s effective number of brands was about twice the length of one answer (2.0 times), ChatGPT’s 1.8 times and Perplexity’s 1.3 times. An answer of six brands from ChatGPT typically comes from a pool of nine or ten that it moves between.

## Five kinds of change

Two answers can differ in which brands they name, in the order, in the first pick, in the sources they cite and in the wording. We measured each separately over every pair of runs of the same question (596 pairs). Each figure below is the mean share that changes between two runs, from 0 (identical) to 1 (nothing shared).

| Change between two runs | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| Brands named | 0.455 | 0.580 | 0.320 |
| Order (rank-biased) | 0.405 | 0.528 | 0.243 |
| Order of shared brands | 0.182 | 0.280 | 0.159 |
| First pick | 0.462 | 0.570 | 0.220 |
| Top 3 | 0.444 | 0.637 | 0.287 |
| Cited domains | 0.525 | 0.674 | 0.013 |
| Cited URLs | 0.637 | 0.779 | 0.020 |
| Words used | 0.588 | 0.689 | 0.551 |

How to read the rows: “brands named” is 1 minus the Jaccard overlap of the two brand lists; “order (rank-biased)” is 1 minus rank-biased overlap, which weighs the top of the list most; “order of shared brands” rescales Kendall’s tau for brands both runs named; “first pick” is the share of run pairs whose first-named brand differs.

Three things stand out. When two runs named the same brands, they mostly kept them in a similar order (the lowest of the brand rows), so most change is in which brands appear, not how they are ranked. The first pick changed about as often as the brand list did. And Perplexity is the outlier on sources: its citations barely moved while its brands and wording still did.

- Chart: [how much changes between two runs, cited URLs vs brands named](https://underneath.agency/research-data/ai-recommendation-consistency-study/sources-vs-brands-change.svg)
- Chart: [share of brands named in every one of five runs](https://underneath.agency/research-data/ai-recommendation-consistency-study/brands-in-every-run.svg)

## The first pick and the top 3

For someone asking once, the brand an answer names first is effectively the recommendation.

| Across 20 questions | ChatGPT | Gemini | Perplexity |
|---|---|---|---|
| First pick changed at least once | 80.0% | 85.0% | 40.0% |
| Mean share of runs with the most common first pick | 71.8% | 63.0% | 87.0% |
| A brand in the top 3 of every run | 60.0% | 40.0% | 90.0% |
| Same top 3 in all runs | 15.0% | 0.0% | 15.0% |

Stable membership did not mean a stable position. Of the brands named in every run, 86.4% moved position at least once on ChatGPT, 87.1% on Gemini and 72.3% on Perplexity; 34.8%, 25.8% and 42.2% of them were in the top 3 every time. A brand can have steady visibility and an unsteady rank.

## Same sources, different brands

If the brands change because the assistant finds different pages each time, runs that cite the same sources should name the same brands. Partly, they do. Within a question, each 0.1 more overlap in cited domains went with 0.03 more brand overlap (95% interval 0.017 to 0.043), with similar slopes for ChatGPT (0.027) and Gemini (0.032).

But Perplexity cited nearly the same sources on every run (mean URL overlap 0.980), and its brands still changed. In the 166 pairs of Perplexity runs with identical cited URLs, the brand lists differed 91.6% of the time (mean overlap 0.675) and the first pick differed 21.1% of the time. The same happened on ChatGPT in the few pairs with identical URLs (9 pairs, 44.4% with a different brand list).

So the evidence an answer shows and the brands it recommends are two different things to measure: evidence can be stable while recommendations are not. Our data measures outputs only, so it cannot say which step inside the assistant produces the change; the cited sources are what the answer shows, not everything the assistant read.

## What makes a question more or less stable

The mean run-to-run overlap ranged from 0.299 for the least stable question to 0.747 for the most stable (averaged over the three assistants). Splitting the variation across the 60 question-assistant cells, the question accounted for 30.0%, the assistant for 25.5% and their combination for 44.5%. That last share matters: a question that was stable on one assistant was not reliably stable on another (correlation 0.301 between ChatGPT and Gemini, around zero between either and Perplexity).

With 20 questions, the patterns by type are exploratory.

- **By industry:** B2B software had the most stable brands (mean overlap 0.708 across assistants), franchises the least (0.434), with retail close behind (0.458).
- **Place-named questions:** on Gemini, the 7 questions naming a city were less stable than the 13 national ones (0.343 vs 0.460); on ChatGPT and Perplexity there was little difference.
- **List length:** on Gemini, questions with longer answers were more stable (correlation 0.573 between brands per answer and overlap); on Perplexity the reverse (−0.431).

## Minutes, hours, rewording and country

The five runs were minutes apart. Our [country study](https://underneath.agency/research/ai-recommendations-by-country-study) and [rewording study](https://underneath.agency/research/ai-prompt-phrasing-study) asked some of the same questions about 4.4 hours later the same day, including a US answer to the unchanged question. That lets us compare, on the same questions, how far an answer moves when you ask again at once, ask again hours later, and change the country or the wording.

| Mean overlap with the five runs | ChatGPT | Gemini |
|---|---|---|
| Another run, same batch (10 questions) | 0.588 | 0.453 |
| Same question, hours later | 0.477 | 0.491 |
| Another country, hours later | 0.398 | 0.363 |
| Another run, same batch (5 questions) | 0.694 | 0.453 |
| Same question, hours later | 0.455 | 0.466 |
| Reworded, hours later | 0.431 | 0.298 |

The first three rows use the 10 questions in the country study; the last three the 5 in the rewording study.

For ChatGPT, time alone moved the answer: the same US question hours later overlapped with the runs less than the runs did with each other. Changing the country moved it further (0.477 to 0.398). Rewording, by contrast, added little beyond time (0.455 to 0.431). An earlier comparison against the minutes-apart runs alone (0.694 to 0.431) would credit that whole gap to the wording.

For Gemini there was no sign of change over the hours (0.453 then 0.491), and both country (0.491 to 0.363) and rewording (0.466 to 0.298) moved the answer. Perplexity was not in the country study and had no hours-later answer; its reworded answers overlapped with the runs 0.462, against 0.663 between runs.

These are single later answers on 5 or 10 questions, so they show the direction, not firm sizes. They do show that short-term repetition, elapsed time and the question itself are separate sources of change, and that a comparison between two conditions needs a baseline collected at the same time.

## What one run gets wrong

A tracker that asks each question once records one draw. Compared with the five runs, a single run did three things.

- It showed 57.8% of the brands the five ChatGPT answers named (95% interval 50.0% to 65.2%), 48.4% for Gemini and 72.2% for Perplexity.
- It left out brands that were not marginal: of the brands missing from a single run, 16.3% on ChatGPT, 15.8% on Gemini and 27.5% on Perplexity were named in at least three of the five.
- It sometimes reversed two brands: for pairs of brands whose five-run frequencies differed by 0.4 or more, a single run named the rarer brand but not the more frequent one 4.2% of the time on ChatGPT, 6.1% on Gemini and 2.0% on Perplexity.

Each extra run kept finding new brands. Averaged over every order of the runs, the second run added 2.3 new brands on ChatGPT, the third 1.4, the fourth 1.1 and the fifth 1.0. On Perplexity the fifth run added 0.3. Five runs do not exhaust the tail on ChatGPT or Gemini.

### How precise five runs are

A frequency from five runs is coarse. Treating runs as independent, the 95% interval for a brand named in 3 of 5 runs is 23.1% to 88.2%; for 5 of 5 it is 56.6% to 100%; for 1 of 5, 3.6% to 62.4%. The median interval in our data spans 0.588. To estimate a frequency near 50% to within 10 points would take about 97 runs, to within 20 points about 25; near 80%, about 62 runs for 10 points. Runs minutes apart may not be independent, and drift over days adds variation, so these are lower bounds on the effort, not a recommendation. Measuring how the estimate settles as runs are added is the next step.

## Observed, inferred and unknown

- **Observed:** brand lists, order, first picks, cited sources and wording across 300 answers, and their change between runs, hours later, across countries and across rewordings on the shared questions.
- **Inferred:** that visibility should be reported as a frequency with an interval; that evidence stability and recommendation stability are distinct; that time and condition effects need time-matched baselines.
- **Unknown:** drift over days and weeks; how many runs a given precision needs in practice; which internal step (retrieval, ranking or writing) produces the change.

## What this means

What follows is our interpretation of the numbers above.

- **Report frequency, not presence.** “Named in 4 of 5 runs” says more than “appeared in ChatGPT”. Give the number of runs and an interval alongside it.
- **Separate membership from position.** A brand can be in every answer and still move between first and sixth. Track the share of runs a brand is named, is in the top 3 and is first, separately.
- **Stable sources do not mean stable recommendations.** Tracking which pages an assistant cites is not a substitute for tracking what it recommends.
- **Compare like with like in time.** A change between two measurements hours or days apart mixes real change with drift and noise; measure the baseline variation at the same time.
- **Aim for the core.** Brands named in every run are the assistant’s default answer for that question. Moving from the tail into the core is a real change; a single appearance is not.

## What comes next

- A longitudinal panel: the same questions on several days over several weeks, to separate minute-scale variation from drift (planned with study 24, on or after 3 October 2026).
- 20 runs on a subset of questions, to measure how the estimated frequency settles as runs are added and how many runs a target precision needs in practice.
- Repeated rewordings: each rewording asked several times, so wording and repetition can be compared on equal terms.
- A fixed-evidence experiment: the same pages given to a model repeatedly, to separate variation in what is found from variation in what is written.

## How this compares with other studies

| Source | Sample and date | Figure |
|---|---|---|
| This study | 20 questions, 5 runs, 3 assistants, runs within minutes, September 2026 | 25.2% of ChatGPT brands in all 5 runs |
| SparkToro and Gumshoe | 2,961 runs of 12 prompts by 600 volunteers, November to December 2025 | Less than a 1 in 100 chance of the same list twice |
| Detailed.com | 70,000+ responses over 28 days, September 2026 | On ChatGPT the core brands changed 13% from one day to the next, the tail 78% |
| Detailed.com | Same | “Only about a quarter of the pages ChatGPT cited for a prompt were cited again the next day” |

The studies agree on the shape: a small, fairly stable core and a long, unstable tail. Ours adds that the variation is there within about a quarter of an hour, that it is mostly in which brands appear rather than their order, and that it happens even when the cited sources are identical. Recent research on measuring AI search visibility makes the same methodological point: visibility should be measured repeatedly and reported as a distribution, and evaluation should look at the full search, ranking and writing pipeline rather than the final answer alone.

Sources: [SparkToro](https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/); [Detailed.com](https://detailed.com/ai-volatility/); [Don’t Measure Once (arXiv)](https://arxiv.org/abs/2604.07585); [SAGEO Arena (arXiv)](https://doi.org/10.48550/arXiv.2602.12187).

## Methodology

- **Questions:** 20 of the 80 questions in the [four-assistant comparison](https://underneath.agency/research/ai-assistants-brand-agreement-study), at least two per industry, 7 naming a place; listed in the dataset.
- **Collection:** five runs per question per assistant on 26 September 2026; the first run is the answer used in the four-assistant study. ChatGPT and Gemini via DataForSEO LLM Scraper (consumer apps, US location); Perplexity sonar via its API with web search. The five runs of a question were a median of about 11 minutes apart (at most 16). No two answers were word-for-word identical.
- **Valid runs:** answers of at least 200 characters. One ChatGPT run (an unfinished reply) is excluded from the frequencies, so that question has four valid ChatGPT runs.
- **Brands:** candidate names from each answer, classified brand or not once per unique name, clustered per question over every answer to it (including the country and rewording answers), then matched by normalized name. Rank is order of first mention.
- **Measures:** frequency with a Wilson interval; core, stable, intermittent and tail classes; entropy and effective number of brands; over all pairs of runs, Jaccard overlap of brands, top 3, cited domains, cited URLs and words, rank-biased overlap (p = 0.9), Kendall’s tau among shared brands, and same first pick.
- **Models:** brand overlap on domain overlap, pooled with assistant indicators and within question and assistant; split of the variation in mean overlap by question and assistant.
- **Intervals:** 95% bootstrap resampling the 20 questions (2,000 resamples, seed 20260926), paired across assistants.
- **Code:** cite/pipeline/s8_v11.py and s8_v11_package.py; version 1.0 figures are kept in stats.json.

## Limitations

- Five runs over about 11 minutes: frequencies are coarse, and drift over days and weeks is not measured.
- 20 questions and 3 assistants; the analysis by question type is exploratory.
- Outputs only: the data cannot show whether the change comes from retrieval, ranking or writing. Cited sources are what the answer shows, not everything the assistant read.
- The hours-later, country and rewording answers are single answers on 5 or 10 questions.
- Perplexity was queried through its API, not the consumer app; Claude was not included.
- Brand matching by name can miss variants; rank is the order of first mention, not a ranking the assistant stated.

## What changed in version 1.1

Version 1.0 (26 September 2026) reported how much the brands change. Version 1.1 (28 September 2026) re-analyzes the same answers as a measurement problem, after an external audit: per-brand frequencies with intervals, core-to-tail classes, concentration, order and first-pick stability, a five-part profile of change, source-brand coupling including runs with identical sources, question characteristics, a time-matched comparison with the country and rewording answers, and what one run gets wrong.

Two method changes moved the headline figures slightly: one unfinished ChatGPT run is now excluded, and the brand list for each question is shared with the country and rewording answers. Brands in all five runs moved from 23.7% to 25.2% (ChatGPT), 13.6% to 13.7% (Gemini) and 41.2% to 40.9% (Perplexity); mean overlap between runs from 0.530 to 0.545, 0.421 to 0.420 and 0.682 to 0.680. The conclusions are unchanged. Version 1.0 figures remain in stats.json.

## Data and downloads

- Every answer with its brands in order, cited domains and condition: [s8_answers_v11.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_answers_v11.csv)
- Every brand’s frequency, interval and ranks per question and assistant: [s8_brand_visibility.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_brand_visibility.csv)
- Every pair of runs with all overlap measures: [s8_run_pairs.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_run_pairs.csv)
- Stability per question and assistant: [s8_question_engine.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s8_question_engine.csv)
- Version 1.0 answer file: [s7_s8_answers.csv](https://underneath.agency/research-data/ai-recommendation-consistency-study/s7_s8_answers.csv)
- Brand classification audit: [brand_classification_audit.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/brand_classification_audit.json)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-recommendation-consistency-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Ask an AI the same question 5 times: do the brands change?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-recommendation-consistency-study

## Frequently asked questions

### Does ChatGPT give the same recommendations every time?

No. Across five runs of the same question minutes apart, 25.2% of the brands ChatGPT named appeared in all five answers and 36.6% in only one. Its first pick changed at least once for 80.0% of questions.

### Which AI assistant gives the most consistent recommendations?

Perplexity, in our test: 40.9% of its brands appeared in all five runs and its first pick stayed the same for 60.0% of questions. Gemini was the least consistent. The ranking describes these 20 questions on one day; how stable a question is also depends on the question itself.

### How many times should you ask an AI assistant to measure brand visibility?

More than once, and more than five times for a precise figure. One ChatGPT answer showed 57.8% of the brands five answers named, and the fifth run still added about one new brand. A brand named in 3 of 5 runs has a 95% interval of 23.1% to 88.2%.

### Why do AI answers change between identical questions?

We measured the outputs, not the mechanism. The change is only partly tied to sources: in 166 pairs of Perplexity runs with identical cited URLs, the brand list still changed 91.6% of the time. Language models generate text with some randomness, and assistants that search the web can also retrieve and rank sources differently, but our data cannot separate those steps.

## Related research

- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)
- [Does rewording a question change AI brand recommendations?](https://underneath.agency/research/ai-prompt-phrasing-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- [ChatGPT local recommendations vs Google Maps](https://underneath.agency/research/chatgpt-local-recommendations-study)

## Related guides

- [Why does ChatGPT give a different answer about my brand each time?](https://underneath.agency/resources/why-ai-answers-about-your-brand-change)
- [Which AI engine gives the most consistent answers about brands?](https://underneath.agency/resources/most-consistent-ai-engine-for-brands)
- [Can I trust a one-time AI visibility report for my brand?](https://underneath.agency/resources/one-time-ai-visibility-report)
- [How many prompts do we need to track to measure AI visibility?](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility)
- [Why do AI search engines cite different sources every time I check?](https://underneath.agency/resources/why-ai-search-citations-change)

---

This is the Markdown twin of https://underneath.agency/research/ai-recommendation-consistency-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
