Guide · AI search

Which AI engine gives the most consistent answers about brands?

No single AI engine is the most consistent on every measure. In the two tests that compared which brands an engine names, Perplexity repeated itself most often; when researchers instead compared how engines describe a brand, Gemini and ChatGPT were steadier and Perplexity least steady. For anyone monitoring a brand, the engine you track shapes how reliable your numbers are, but so does the question you ask.

The short version

  1. In our test of 20 buyer questions asked five times each, 40.9% of the brands Perplexity named appeared in every run, against 25.2% for ChatGPT and 13.7% for Gemini.
  2. A Swiss study of four engines found the same order for brand lists repeated within 24 hours: Perplexity overlapped most (0.492), Google’s AI Mode least (0.375).
  3. On shopping questions asked from the Netherlands, Google’s AI Overviews repeated their product picks most (0.421 overlap) and ChatGPT least (0.178).
  4. When a vendor compared the wording of repeated answers about 20 European brands, Gemini (0.952) and OpenAI’s model (0.950) were steadier than Perplexity (0.904).
  5. In our test, the question explained more of the variation in consistency (30.0%) than the choice of assistant (25.5%).

Which engine repeats the same brands most often?

Perplexity, in both independent tests that included it, with ChatGPT in the middle and Gemini or Google AI Mode last.

In our consistency study, we asked ChatGPT, Gemini and Perplexity the same 20 US buyer questions five times each on one day in September 2026. Perplexity named 40.9% of its brands in all five runs. ChatGPT managed 25.2% and Gemini 13.7%. The first brand named changed at least once for 80.0% of questions on ChatGPT, 85.0% on Gemini and 40.0% on Perplexity.

Schulte and colleagues (opens in a new tab) ran German-language prompts in four Swiss product categories up to ten times within a day, in March 2026. Measured as the share of brands two runs had in common, Perplexity scored 0.492, ChatGPT 0.437, Gemini 0.409 and Google AI Mode 0.375. One author also lists an industry affiliation, and all data came from Swiss servers.

Measure of brand-list repeatabilityMost repeatableLeast repeatable
Brands named in all 5 runs (our study, US)Perplexity 40.9%Gemini 13.7%
Brands shared between two runs (Swiss study)Perplexity 0.492Google AI Mode 0.375
Products shared between repeats (Dutch audit)AI Overviews 0.421ChatGPT 0.178

Does the ranking hold for product picks and cited sources?

Only partly: the order changes with the task, the country and whether you look at brands or sources.

Uberti-Bona Marin and colleagues (opens in a new tab), independent academics, asked 117 real shopping questions three times each from the Netherlands in September 2026. They did not test Perplexity. Google’s AI Overviews repeated the most products between runs (0.421), Gemini 0.287 and ChatGPT only 0.178. The sources moved too: repeated answers shared 26.0% of cited websites on ChatGPT, 29.8% on Gemini and 45.9% on AI Overviews.

Sources and brands can also tell different stories. In the Swiss study, Gemini had the most stable sources, though not the most stable brands. In a June 2026 test of 15 prompts by Tannenbaum (opens in a new tab), founder of a visibility software firm, 82.6% of ChatGPT’s cited pages turned over from one day to the next, against 45.5% on Perplexity. He notes that collection changes may explain part of that. Sielinski (opens in a new tab), from another vendor, found OpenAI’s search often either repeated its sources exactly or changed them completely. For the reasons behind that churn, see why cited sources change between checks.

Which engine describes a brand most consistently?

Gemini and ChatGPT, in the one study that measured how similar the wording of repeated answers was.

Żatuchin (opens in a new tab), who works for an AI brand-monitoring company, asked three engines about 20 Central and Eastern European brands, repeating each prompt five times in spring 2026. Repeated answers were very similar overall: average similarity was 0.935 on a scale up to 1, and 87.2% of comparisons scored above 0.90. By engine, Gemini (0.952) and OpenAI’s model (0.950) were steadier than Perplexity (0.904).

This does not contradict the brand-list results. Perplexity can name the same brands while phrasing its answer differently each time. Which engine looks “most consistent” depends on whether you track names or narrative.

Does the engine matter more than the question or the language?

Not always: the question, and for some measures the language, matter as much as the engine.

In our study, the question accounted for 30.0% of the variation in stability and the assistant for 25.5%. A question that was stable on one assistant was not reliably stable on another. Żatuchin found little difference between languages in how steady the wording was: English scored 0.941 and Estonian 0.925.

A second paper by Żatuchin (opens in a new tab), scoring how positive answers were about the same brands, found language mattered far more. Query language explained 26.5% of the variation in a single answer, the brand itself only 1.5%. That measure was weak, though: 91.9% of answers scored exactly neutral. Its practical point stands: a single answer says almost nothing about where a brand stands. The same caution applies to a one-off AI visibility report.

Does a consistent engine give more trustworthy answers about you?

No: consistency only means the answer repeats, not that it is right or complete.

Perplexity’s steadiness shows the gap. In our study, 166 pairs of Perplexity runs cited exactly the same pages, yet the brand list still differed 91.6% of the time. Stable sources did not guarantee stable recommendations.

A vendor dataset from Kumar at Ranqo (opens in a new tab), covering 102 brands, found brand mentions mostly fixed: only 6.8% of brand-question-engine combinations flipped between being named and not. But the tone of those mentions flipped 45.5% of the time. An engine can be reliable about whether it names you and unreliable about how it talks about you.

What should you do about it?

Choose engines by where your buyers ask, then measure each one with enough repeats to see its real pattern.

  1. Do not pick an engine to monitor because it looks stable. Track the engines your customers use, not just one.
  2. Run each question several times per engine. One ChatGPT answer in our study showed only 57.8% of the brands its five answers named.
  3. Track separately whether you are named, where you rank and how you are described. Each moves at a different rate.
  4. Expect more noise from ChatGPT and Gemini brand lists, and set wider thresholds before reacting to a change there.
  5. Test your questions in every language your buyers use, since language can matter as much as engine.

If you want this set up as an ongoing program, see our generative engine optimization service.

What does the research not tell us yet?

It does not tell us whether these rankings hold over weeks, in other categories or after engine updates.

  • Every study is a snapshot of one to five days; engines change models often.
  • Samples are small: 20 questions in our study, 15 prompts in Tannenbaum’s, 20 brands in Żatuchin’s.
  • Three of the studies are by vendors that sell visibility measurement.
  • Claude and Microsoft Copilot are missing from most comparisons.
  • No study links an engine’s consistency to whether buyers trust or act on its answers.

Frequently asked questions

Is Perplexity more consistent than ChatGPT?

For brand lists, yes, in two tests. In ours, 40.9% of Perplexity’s brands appeared in all five runs against 25.2% for ChatGPT; for how answers are worded, one study found Perplexity less steady.

Why does ChatGPT give different brand recommendations each time?

AI assistants generate each answer afresh and may search the web differently each run. In our test, ChatGPT’s first-named brand changed at least once for 80.0% of questions.

How many times should we run a prompt to measure AI visibility?

More than once, and more on less stable engines. A single ChatGPT answer showed 57.8% of the brands its five runs named between them.

Does the language of the question affect AI answers about brands?

It can matter more than the engine for some measures. In one study, query language explained 26.5% of the variation in tone of a single answer.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.