Guide · AI search

Why does ChatGPT give a different answer about my brand each time?

Because ChatGPT writes each answer by sampling, so asking the same question twice can produce different brands, sources and wording. In one detailed study, about a third of the movement between identical questions was this built-in chance alone. A single answer is one draw from a range of possible answers, so it should never drive a decision on its own.

The short version

  1. In one study, chance alone accounted for 34.8% of the variation between answers to identical questions, even at a low randomness setting (Żatuchin, 2026 (opens in a new tab)).
  2. In our test of 20 buyer questions asked five times, only 25.2% of the brands ChatGPT named appeared in every run, and 36.6% appeared once.
  3. In a Dutch audit, ChatGPT’s product picks overlapped by 0.178 between repeats of the same question, against 0.421 for Google’s AI Overviews (Uberti-Bona Marin and colleagues, 2026 (opens in a new tab)).
  4. ChatGPT’s first-named brand changed at least once for 80.0% of our questions over five runs.

Is ChatGPT broken when its answer changes?

No; varying answers are how these systems work. An AI assistant builds each answer word by word, choosing among likely options with some randomness. With web search switched on, it may also run different searches and read different pages each time.

Żatuchin (opens in a new tab) measured how much of the movement comes from this chance alone. Three AI models answered the same questions about 20 Central and Eastern European brands about five times each. When everything else was held fixed, chance accounted for 34.8% of the variation in the answers’ scores. The models ran at a low randomness setting, so the author treats that figure as a floor; default app settings are likely to vary more.

The author is affiliated with Rankfor.AI, which sells brand monitoring, and the score measured was the tone of each answer. Even so, the finding matches every other study below: identical questions do not give identical answers.

How much do the answers actually change?

Enough that one answer shows only about half of what the assistant tends to say. In our consistency study, we asked ChatGPT, Gemini and Perplexity 20 buyer questions five times each, with the runs a median of about 11 minutes apart.

Brands named across five runsChatGPTGeminiPerplexity
Named in all five25.2%13.7%40.9%
Named in only one36.6%46.7%15.8%
Share of five-run brands seen in a single run57.8%48.4%72.2%

So most brands sit in a long tail that comes and goes. A small core appears nearly every time.

An independent audit found the same pattern for shopping questions. Uberti-Bona Marin and colleagues (opens in a new tab) sent 117 real product questions to the ChatGPT and Gemini apps and to Google, three times each, from the Netherlands. ChatGPT’s recommended products overlapped by only 0.178 between repeats, on a scale where 1 means identical. That was about as different as ChatGPT was from Google’s answers.

Is it chance, or did something change on the web?

Mostly chance; the pages behind the answers rarely change that fast. Sielinski (opens in a new tab) stored a fingerprint of each cited page he could reach during nine days of repeated questions to three AI search engines. Most pages did not change at all between collections, yet citations still swung.

His conclusion was that swings of this size, including rank changes within a single four-hour window, are built into the engines rather than caused by edits to the pages. On daily data, his uncertainty ranges for a site’s share of citations were 3 to 7 percentage points wide.

A Swiss study pointed the same way. When the same prompts were asked within 24 hours, cited sources overlapped by only 0.32 to 0.43, about the same as from one day to the next (Schulte, Bleeker and Kaufmann, 2026 (opens in a new tab)). In our study, Perplexity cited identical pages in 166 pairs of runs, and the brand list still changed in 91.6% of them.

Which parts of an answer change most?

Rankings and wording move more than the basic fact of whether you are mentioned. In our study, ChatGPT’s first-named brand changed at least once for 80.0% of questions over five runs. Even among brands named every time, most moved position.

The framing shifts too. In the Dutch audit, whether an assistant labeled a product “the best” changed across repeats for between 27% and 40% of questions, depending on the assistant. Whether an answer named any product at all changed far less often.

A study by Ranqo, a company that sells AI visibility tracking, found the same split across more than 100 brands. For brand, question and engine combinations tracked over several runs, the tone flipped between positive and negative in 45.5% of them. Whether the brand was mentioned flipped in only 6.8% (Kumar, 2026 (opens in a new tab)). Most of those combinations were small brands that were never mentioned, which keeps the second figure low. Still, tone is the noisiest thing you can track.

Do other AI assistants vary as much as ChatGPT?

No; each assistant has its own level of consistency. In our study, Perplexity was the steadiest, naming 40.9% of its brands in every run. Gemini was the least steady at 13.7%. In the Dutch audit, Google’s AI Overviews were the most consistent, with product picks overlapping by 0.421 between repeats. Which engine looks steadiest depends on the measure; see which AI engine is most consistent.

Repeats also keep turning up new sources. In the Dutch audit, the number of distinct websites ChatGPT showed for a question rose from 2.65 after one request to 5.56 after three. Even Google varies in whether it shows an AI answer at all: an earlier audit cited by Sielinski found a 33% inconsistency rate in whether AI Overviews appeared for the same query.

A survey of the field adds that randomness is not the whole story. It cites an audit in which repeated runs with the randomness setting at zero still changed 9 to 28% of decisions (Martinez, 2026 (opens in a new tab)).

What should you do about it?

Treat one AI answer as one roll of the dice, and judge your brand over many runs. In practice:

  1. Never act on a screenshot. A missing mention in one answer, or a sudden first place, is weak evidence on its own.
  2. Ask each important question many times and report a rate, such as “named in 7 of 10 runs”, with the number of runs beside it.
  3. Track separately whether you are named, whether you are near the top and how you are described. These move at very different speeds.
  4. Aim to be in the core: the brands an assistant names nearly every time. Moving out of the long tail is a real gain; a single appearance is not.
  5. Compare assistants separately. A brand can be steady on Perplexity and erratic on Gemini.
  6. When a number moves, check it against the normal run-to-run range before you treat it as news. Our guide on whether a GEO campaign really worked shows how.

If you want a tracking setup built on repeated runs, see our generative engine optimization service.

What does the research not tell us yet?

We know the answers vary; we know less about how much is enough to measure. Gaps:

  • No study yet separates exactly how much variation comes from the search step and how much from the writing step inside the assistant.
  • Most repeat tests are short: minutes to nine days. How a brand’s steady core changes over months, through model updates, is barely measured.
  • Several studies here come from companies that sell AI visibility tools. Their data are useful, but independent checks are fewer.
  • The Dutch audit used three repeats per question; our study used five. Both say these are too few to capture every brand an assistant might name.
  • Logged-in users with chat history may see different answers again. The studies here used logged-out sessions or developer access.

Frequently asked questions

Why does ChatGPT recommend different brands every time I ask?

Because it generates each answer with some randomness and may search different pages each time. In our five-run test, only 25.2% of the brands ChatGPT named for a question appeared in all five answers.

Is one ChatGPT answer a reliable measure of my brand’s visibility?

No. In our test a single ChatGPT answer showed only 57.8% of the brands its five answers named between them, so one answer misses a large share of the picture.

Does Google’s AI Overview change as much as ChatGPT?

Less, in the evidence so far. In a Dutch audit of product questions, AI Overview picks overlapped by 0.421 between repeats, against 0.178 for ChatGPT.

If the AI answer about my brand changed, did a website change?

Usually not. One study that fingerprinted the cited pages found most pages unchanged while citations still swung, which points to the engine rather than the web.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.