Research · AI assistants

Does rewording a question change AI brand recommendations?

Buyers with the same need ask in different words. One says they run a small business, one says money is tight, one asks for a straight answer. This study asks how stable the brands ChatGPT, Gemini and Perplexity recommend are when the question is reworded, and how much of the change comes from the wording rather than from the ordinary variation between two runs of the same question. On 28 September 2026 we asked 20 buyer questions in four wordings, five times each, on all three engines, and three more phrasings of the budget request: 1,560 answers. Every rewording moved the recommendations further than a plain rerun did. A budget moved them most, and an explicit request for an unbiased answer moved them least. That other phrasings change AI recommendations is already established (arXiv 2605.27440 (opens in a new tab)). What this study adds is a same-batch rerun baseline for every question, a split of the change into its parts, and a test of whether the budget effect comes from the constraint or from the particular words.

The short version

  1. Asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it 15.3%, adding “I run a small business with about 10 employees” 40.6%, and asking for “an honest, unbiased answer” 54.9%.
  2. Every rewording changed the brand list by more than a rerun. Brand overlap with the original (Jaccard) was 0.638 for a rerun, 0.547 for the unbiased request, 0.389 for business context and 0.326 for a budget. The excess over rerun variation was 0.273 for a budget (95% interval 0.220 to 0.322) and 0.058 for the unbiased request (0.036 to 0.080).
  3. The wording explained 46.7% of the variation in brand lists between the budget and original answers to the same question, and 21.6% for the unbiased request; the rest is run-to-run variation. By chance alone it would explain about 0.11.
  4. The budget effect comes mainly from the constraint, not the exact words. Four different ways of saying “cheap” overlapped with each other at 0.474, far closer to a same-phrasing rerun (0.563) than to the original question (0.341). The words mattered more on Perplexity than on ChatGPT or Gemini.
  5. The top brand is not protected. A brand named first in the original survived a rerun 94.3% of the time but a budget rewording only 54.5%. Brands that came in on a budget were given a price or value reason 55.9% of the time (model-coded).

Research questions

QuestionAnswered here?
RQ1. Does adding context, a constraint or a meta-instruction change the brands recommended?Yes
RQ2. Is that change larger than rerunning the identical question?Yes, for every question, engine and wording
RQ3. How do the lists change: membership, first brand, top three, order, length?Yes
RQ4. Does sensitivity differ by engine and category?Yes, with a mixed model
RQ5. Which brands hold their place and which appear only under one wording?Yes
RQ6. Does the budget effect come from the constraint or the words?Yes, with four phrasings

Hypotheses and results

HypothesisResult
H1. Adding decision information lowers overlap with the originalSupported
H2. A budget changes the list more than an unbiased-answer requestSupported
H3. Rewordings change the list more than a rerunSupported for all three
H4. Sensitivity differs by enginePartly
H5. Some brands are far more wording-sensitive than othersSupported
H6. Rewording removes lower-ranked brands more than top onesNot for a budget

H1: business context (0.389) and budget (0.326) against a rerun (0.638). H2: the intervals for budget (0.272 to 0.386) and the unbiased request (0.503 to 0.590) do not overlap. H3: the excess over reruns is above zero for all three rewordings, though small for the unbiased request. H4: the raw amount of change did not differ by engine once wording was accounted for (p = 0.304), but measured against each engine’s own rerun variation Perplexity was the most wording-sensitive (p = 0.006). H5 and H6 are covered under the brand findings below.

A model of what moves a recommendation

We treat an AI recommendation as the result of several inputs: the buyer’s underlying need (the question), the wording, the engine, the date and the run. This study changes the wording while holding the question, engine and date fixed, and measures run-to-run variation directly by asking each wording five times. A brand therefore does not have one AI visibility score. It has a visibility for a question, in a wording, on an engine, at a time.

The three rewordings are different kinds of change, not three paraphrases.

RewordingKind of changeExample
Business contextWho is asking“I run a small business with about 10 employees. What is the best air fryer?”
BudgetA constraint“What is the best air fryer on a tight budget?”
Unbiased answerHow to answer“I’m comparing options and want an honest, unbiased answer, not a sales pitch: what is the best air fryer?”

What we measured

We took the 20 national buyer questions used in the first version of this study, drawn from our four-assistant study, such as “What is the best air fryer?” and “Which email marketing platform is best for ecommerce?”.

  • Main design (1,200 answers): each question in the original wording and the three rewordings, five runs each, on ChatGPT and Gemini (their consumer apps) and Perplexity (sonar with web search), United States, 28 September 2026.
  • Budget phrasings (360 answers): the budget request stated three other ways, two runs each on each engine: “What is the best affordable air fryer?”, “What is the best air fryer for someone trying to spend less?” and “Keeping the total price as low as possible is my priority. What is the best air fryer?”.
  • Same batch: the answers to all wordings of a question on one engine were collected side by side, a median of 0.1 minutes apart. The whole collection took 0.4 hours, so time cannot explain the differences between wordings.

Brands were identified as in the four-assistant study: names in bold or headings, classified as brand or not, name variants merged per question, then matched in each answer in order of first mention. For each pair of answers we measured brand overlap (Jaccard: brands in both divided by brands in either), whether the first brand matched, how many of the first three were shared, rank-biased overlap (an agreement score that weights the top of the list most), the share of the original’s brands kept, the share of new brands, the order of shared brands (Kendall’s tau) and the change in list length. The rerun baseline compares the five runs of the same wording with each other. Every 95% interval comes from resampling the 20 questions, because answers to one question are not independent.

Findings

Every rewording moved the list further than a rerun

Against the originalOverlapFirst brandExcessInterval
Same wording, run again0.63868.0%––
Unbiased answer0.54754.9%0.0580.036 to 0.080
Business context0.38940.6%0.1940.154 to 0.240
Budget0.32615.3%0.2730.220 to 0.322

Overlap is brand overlap (Jaccard), first brand the share with the same first brand, and interval the 95% interval of the excess. Excess change is the drop in brand overlap beyond what reruns of the two wordings show on their own. The excess was positive in 93.3% of the 60 question-and-engine pairs for a budget, 91.7% for business context and 76.7% for the unbiased request. For the first brand, the budget excess was 48.8 percentage points.

How much of the variation is the wording

Within each question and engine we split the variation in brand lists into a part explained by the wording and a part left over from run to run (a permutation analysis of variance on brand-list distances).

WordingExplainedSignificant
Unbiased answer21.6%28.3%
Business context38.6%68.3%
Budget46.7%81.7%

Explained is the share of variation explained by the wording; significant is the share of the 60 question-and-engine pairs where that effect is significant at the 5% level. With two wordings and ten answers, chance alone would give about 0.11. Taking all four wordings together, the wording explained 50.4% of the variation (95% interval 46.6% to 53.7%) and was significant at the 5% level in 95.0% of question-and-engine pairs.

A mixed model of the change per question, engine and rewording, with a random effect for each question, confirms the ordering: relative to the unbiased request, a budget raised the set-change score by 0.221 (0.164 to 0.277) and business context by 0.158 (0.102 to 0.214), and the wording as a whole was highly significant (p < 0.001). Differences between questions accounted for about a quarter of the remaining variance (0.255).

Engines differ in noise more than in sensitivity

Brand overlap with the originalRerunUnbiasedContextBudget
ChatGPT0.6420.5550.4080.363
Gemini0.5450.4690.3270.305
Perplexity0.7270.6160.4310.310
Same first brand as the originalRerunUnbiasedContextBudget
ChatGPT74.5%57.4%48.5%21.8%
Gemini50.5%40.4%33.3%13.6%
Perplexity79.0%66.8%40.0%10.4%

Gemini’s lists change the most in absolute terms, but much of that is its own run-to-run variation: two runs of the identical question shared only 0.545 of their brands. Perplexity is the most repeatable (0.727) and so the most wording-sensitive once its low rerun variation is taken into account: the budget excess was 0.414 on Perplexity against 0.226 on ChatGPT and 0.179 on Gemini, and in the model of excess change Perplexity added 0.113 (0.069 to 0.157). The raw amount of change did not depend on the engine and wording together (p = 0.304).

How the lists change

Measure (all engines)RerunUnbiasedContextBudget
Brand overlap0.6380.5470.3890.326
Top three shared75.5%66.7%51.6%33.1%
Rank-biased overlap0.7940.7130.5540.418
Original brands kept–72.0%49.2%43.7%
New brands–31.0%39.3%48.2%
Order of shared brands–0.4950.3610.100
Change in list length–+0.24−1.00−1.12

Rank-biased overlap and order are scores from 0 to 1, where 1 is identical. A budget replaced about half the list and scrambled the order of the brands that stayed: the order score of 0.100 (interval −0.03 to 0.235) is not distinguishable from random. Business context and a budget both shortened the list by about one brand; the unbiased request slightly lengthened it.

The top brand is not protected

For each brand in an original answer we asked how often it appeared in a rerun, and how often in a reworded answer.

Place in the originalRerunUnbiasedContextBudget
First94.3%87.6%69.1%54.5%
Second or third85.0%80.3%58.4%42.8%
Fourth or lower65.4%58.5%35.7%34.2%

Based on 293 first-placed, 574 second- or third-placed and 903 lower-placed brand mentions. Lower-placed brands are less stable in every condition, but a budget costs the first brand about as much as the rest of the list (it keeps 54.5% of its 94.3%). We expected rewording to trim mainly the tail of the list (H6); for a budget it does not. Business context does hit the tail harder.

Which brands hold their place

For each brand, question and engine we counted how often the brand appeared in the five runs of each wording (1,074 brand-question-engine combinations).

Brand patternShare
Named in only one or two of about 20 answers40.6%
Leaves when a budget is added15.0%
No clear pattern14.5%
Enters when a budget is added9.4%
Stable: in at least 60% of runs of every wording7.4%
Enters with business context5.0%
Leaves with business context4.9%
Moves with the unbiased request3.2%

Of the 337 brands named in at least 60% of the original’s runs, 78.6% held that level under the unbiased request, 51.3% under business context and 42.7% under a budget, and 47.8% were named in at most one of the five budget runs. Among the 638 brands named in three or more answers, 88.6% had a visibility that differed by 40 points or more between their best and worst wording. Engine-specific brands, named in half of one engine’s answers and almost never by the other two, were 10.5% of the 143 brands that reached that level anywhere.

A brand does not have one AI visibility score

Perplexity, “Which email marketing platform is best for ecommerce?”, runs out of five that named each brand.

BrandOriginalContextBudgetUnbiased
Omnisend5555
Klaviyo5505
Mailchimp5502
Drip5002
ActiveCampaign4005
MailerLite0354

Klaviyo was named in every Perplexity answer except the budget ones. On the same budget question ChatGPT named Klaviyo in all five runs and Gemini in one. Asked for the best air fryer, Perplexity named Cosori, TurboBlaze, Ninja and Instant Vortex Plus in all five original runs; on a tight budget it led with Gourmia and the Chefman TurboFry 2 Quart, and named Ninja in none. A single “AI visibility” figure for any of these brands would average over wordings that give 0 and 5 out of five.

The constraint matters more than the words

To test whether the budget effect comes from the constraint or from the phrase “tight budget”, we asked the same constraint four ways.

Comparison (all engines)Brand overlapSame first brand
Original wording, run again0.63868.0%
Same budget phrasing, run again0.56360.1%
Two different budget phrasings0.47448.6%
A budget phrasing and the original0.341–

The phrasings agreed with each other much more than any of them agreed with the original: overlap with the original was 0.326 for “tight budget”, 0.370 for “affordable”, 0.353 for “someone trying to spend less” and 0.338 for “lowest total price is my priority”. Brands that came in with “tight budget” also appeared with each other phrasing 69.6% of the time. The words still mattered a little. The gap between a same-phrasing rerun and a different phrasing was 0.090 across engines (0.061 to 0.117), but it was small on ChatGPT (0.031) and Gemini (0.042) and larger on Perplexity (0.196), whose answers depend more on the exact search terms. A recent study of API models found that the prompt string, more than the buyer’s intent, decided which brands surfaced. In these consumer assistants, the intent (a price limit) did most of the work.

What changes with the wording: sources and prices

Each engine cites web pages with its answers, so we can see whether a reworded question also draws on different sources.

DomainsCitesRerunUnbiasedContextBudget
ChatGPT97.5%0.6320.3900.2920.282
Gemini95.2%0.3250.1900.1070.147
Perplexity100.0%0.9860.8200.5170.595

Overlap of cited domains with the original; “cites” is the share of answers citing any source. Rewording changed the cited sources far more than rerunning did, and the questions whose sources changed most were the ones whose brands changed most (Spearman correlation 0.35 on ChatGPT, 0.344 on Gemini and 0.427 on Perplexity, each p < 0.01). Perplexity cited almost the same sites on every rerun (0.986), so its brand changes follow the changed search. The answers also changed in content: 84.7% of budget answers quoted a dollar figure, against 33.0% of original answers, 37.3% with business context and 43.0% with the unbiased request.

Why brands came in with a budget

For the 10 questions where a budget changed the list most, we took each brand that appeared in the first budget answer but not in the first original answer (102 brands) and coded the main reason the answer itself gave for it. Two independent Claude model coders agreed on 87.3% (Cohen’s kappa 0.799); a third model pass settled the 13 disagreements. These are model-coded, not human-validated.

Reason given for a new brandShare
Price or value55.9%
Features21.6%
Suited to a type of buyer7.8%
Quality or reliability6.9%
Named as an alternative, no reason5.9%
No reason2.0%

Price or value was the main reason for 71.0% of new brands on Perplexity, 55.9% on ChatGPT and 43.2% on Gemini. The change was mostly explained, not arbitrary: the assistants brought in brands they could justify on price.

Where rewording matters most

CategoryQuestionsRerunContextBudget
Consumer products50.5550.3860.194
Business software50.7080.5260.363
Insurance and finance70.6100.2890.341
Health and legal services30.7250.3980.451

Brand overlap with the original. Consumer products were the most budget-sensitive: robot vacuums (a set-change score of 0.844, where 1 means no brand in common), noise-cancelling headphones (0.828) and air fryers (0.827). The least sensitive were online LLC formation (0.355), car insurance for young drivers (0.486), whose original question already asks for the cheapest, and robo-advisors (0.502). Questions with more stable reruns were somewhat less budget-sensitive (Spearman 0.29, p = 0.025). The category groups are small, so these differences are descriptive.

Two days apart, and the first version

Answers to the original wording on 26 September (from the first version of this study) overlapped with the 28 September runs about as much as two same-day runs did: 0.642 against 0.642 on ChatGPT, 0.546 against 0.545 on Gemini, and 0.688 against 0.727 on Perplexity. Over two days, time added little beyond run-to-run variation.

Average of three enginesv1.0v1.1
Same first brand, budget12.0%15.3%
Same first brand, business context40.0%40.6%
Same first brand, unbiased55.0%54.9%
Brand overlap, budget0.3660.326
Brand overlap, business context0.4120.389
Brand overlap, unbiased0.5740.547

Version 1.0 used one run per wording on 26 September; version 1.1 five runs per wording on 28 September. The first version’s ordering holds on all three engines. One difference: in version 1.0, ChatGPT’s budget answers overlapped with the original more than its business-context answers did (0.454 against 0.402). With five runs, the budget is the largest change on every engine.

A prompt sensitivity index

Because no single number captures how a list changes, we report sensitivity as three scores, each from 0 (no change) to 1 (complete change), between the original and the reworded answers.

WordingSet changeFirst-brand changeRank change
Same wording, run again0.3620.320.206
Unbiased answer0.4530.4510.287
Business context0.6110.5940.446
Budget0.6740.8470.582

Set change is 1 minus brand overlap, first-brand change is 1 minus the same-first-brand rate, and rank change is 1 minus rank-biased overlap. The first row is the floor that ordinary variation sets. A tracker that reports one prompt run once cannot tell a wording effect from that floor.

How this compares with other studies

StudyWhat changedEnginesRunsMain figure
Aggarwal et al. 2024 (GEO)The source pagesGenerative enginesControlledPage edits change how visible a source is
Chen et al. 2025Engine, language, vertical, paraphraseSeveral AI search enginesControlledCited sources differ by engine and phrasing
Jack et al. 2026, personaA persona prefixOpenAI, Anthropic APIs2,000Jaccard lowered by 0.12 to 0.20
Jack et al. 2026, paraphraseCosmetic and constraint paraphrasesOpenAI, Anthropic APIsAbout 12,0000.288 and 0.135, against 0.50 to 0.61 for reruns
Rankshift 2026Seven near-synonym CRM promptsSeveral1,176 eachHubSpot at 98% on every variant
Tannenbaum 2026None (monitoring data)GPT, Gemini34,960Mentions depend on the brand being retrieved
This studyContext, constraint, meta-instruction, four budget phrasingsChatGPT, Gemini, Perplexity apps5 per wording0.326 for a budget against 0.638 for reruns

The API study of paraphrases found constraint-adding rewordings overlapping at 0.135 against 0.50 to 0.61 for reruns. Our budget rewording (0.326 against 0.638) points the same way, less extremely. The clearest difference is on cosmetic paraphrase: there, different wordings of the same intent overlapped at 0.288; here, different wordings of the same budget constraint overlapped at 0.474, close to a rerun on ChatGPT and Gemini. Persona conditioning lowered overlap by 0.12 to 0.20, and our business-context wording, a short persona, lowered it by more (0.194 beyond reruns). A study of monitoring data found that brands are mentioned mainly when their own pages appear in what the engine retrieves. Our finding that brand changes follow source changes is consistent with that.

Sources: arXiv 2311.09735 (opens in a new tab); arXiv 2509.08919 (opens in a new tab); arXiv 2605.30207 (opens in a new tab); arXiv 2605.27440 (opens in a new tab); Rankshift (opens in a new tab); arXiv 2609.23162 (opens in a new tab).

Observed, inferred and unknown

What we observe. Asked the same question in the same batch, the three assistants return lists that vary from run to run, and adding context or a constraint changes the list well beyond that. A budget changes the first brand in most answers, replaces about half the brands and scrambles the order of the rest. Several ways of stating a budget lead to similar lists. The cited sources change along with the brands, and the new brands come with price reasons.

What we infer. The assistants treat a budget or a stated situation as a different request and search for it differently, rather than reacting to particular words. AI brand visibility is conditional on how the need is expressed, so it should be measured across wordings and repeated runs, not from one prompt run once.

What remains unknown. Whether the new recommendations are better or worse for the buyer: we measured difference, not quality. How much answers drift over weeks rather than two days. Whether logged-in history or personalization adds more variation. Which brand characteristics (price level, size, review volume) predict holding a place across wordings: the brand-level data is published for that analysis, but we have not collected those attributes.

What this means

The points below are interpretation. They follow from the findings but were not tested.

  • Report visibility for a question in a wording, not one score. A brand that appears in every answer to “What is the best email marketing platform for ecommerce?” can be absent from every answer that adds a budget. Measure each important wording separately.
  • Run each prompt several times before reading a change. Two runs of the identical question shared 0.638 of their brands. A difference smaller than that is often noise.
  • Match the buyer’s constraint, not their exact phrase. Different ways of asking for something cheap led to similar lists, so content that states prices and value plainly serves all of them.
  • Being first in the default answer does not carry over. The first-named brand kept its place in only about half of budget answers.

Where this sits in GEO research

Generative engine optimization began by asking whether changing a source page changes how visible it is in AI answers. Later work showed that AI search engines differ by language, vertical and phrasing. A 2026 survey of the field names repeated runs, paraphrase, denominators and causal measurement as unsolved measurement problems. This study takes up two of them: it separates wording effects from rerun variation with a same-batch baseline, and it shows that a brand’s visibility has no meaning without saying for which wording. It does not show how to change that visibility. The next collection repeats the main design on a new date to separate time from run-to-run variation.

Methodology

  • Questions: 20 national buyer questions from the four-assistant study (every second national question), each in the original wording and three rewordings; three further budget phrasings.
  • Engines: ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper; Perplexity (sonar) via DataForSEO LLM Responses API with web search; United States; English.
  • Collection: 1,560 requests on 28 September 2026 between 08:57 and 09:19 UTC: 1,200 in the main design (five runs per wording) and 360 for the budget phrasings (two runs each). All 1,560 returned; six unfinished ChatGPT answers under 200 characters are kept as outcomes but excluded from brand measures, and 18 answers named no brand. Version 1.0 answers (26 September 2026, one run per wording) are used for the cross-date and version comparisons.
  • Brands: candidate names from bold text and headings (and ChatGPT’s brand panels); each unique string classified as brand or not (study 7 decisions reused; new strings first filtered by rule for prices, pronouns and long phrases, then classified by Claude model coders, with directive fragments and URLs corrected by rule); name variants merged per question; matched in each answer in order of first mention. Names that are also ordinary words (such as Choice, Fidelity or Travelers) count only where the answer lists them as an item, and review sites cited as sources are not counted as brands.
  • Measures: Jaccard similarity, same first brand, top-three overlap, extrapolated rank-biased overlap (persistence 0.9), retention, replacement, Kendall’s tau on shared brands (three or more), list length. Rerun baseline from all pairs of the five runs; between-wording figures from all 25 pairs of original and reworded runs.
  • Variance and models: permutation analysis of variance on Jaccard distance per question and engine (499 permutations); linear mixed model with wording and engine as fixed effects and question as a random intercept; likelihood-ratio tests.
  • Reasons for change: 102 brands entering in the first budget answer against the first original answer, for the 10 most budget-sensitive questions; two independent model coders, third-pass adjudication.
  • Uncertainty: 95% bootstrap intervals resampling the 20 questions (2,000 resamples, seed 20260928).
  • Update schedule: quarterly; next collection repeats the main design on a new date.

Limitations

  • Three engines, 20 questions, the United States and English, with the main design on one day. Other categories, countries and dates may differ.
  • ChatGPT and Gemini answers come from their consumer apps through a data provider, without a signed-in history; Perplexity is the API with web search. Personalization is not measured.
  • Four wordings and four budget phrasings are a small sample of how buyers ask. The business-context wording was added to every question, including consumer ones where it is less natural.
  • Five runs per wording estimate run-to-run variation with some noise, and two runs per budget phrasing is thin.
  • Brand identification is automatic and can merge or split names; the reasons for change are model-coded, not human-validated.
  • We change the question, not the model or its retrieval, so the source analysis shows association, not cause.
  • Different recommendations are not necessarily better or worse ones.

What changed in version 1.1

Version 1.0 (26 September 2026) asked each rewording once and compared it with the original from our four-assistant study, with a rerun comparison on only 5 questions from a batch collected hours earlier. Following an external review, version 1.1 (28 September 2026) is a new collection of 1,560 answers. It adds five same-batch runs of every wording for all 20 questions, measures of how lists change (not only Jaccard), a variance split and mixed model with intervals clustered by question, brand-level patterns, model-coded reasons for change, a test of four budget phrasings, hypotheses taken from the review before the new data were collected, and a comparison with the literature. Brand matching now ignores ordinary words and review sites that the earlier method counted as brands. The version 1.0 figures are kept in the data files and compared above; its conclusions hold.

Data and downloads

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). Does rewording a question change AI brand recommendations? (Version 1.1). Underneath Research. https://underneath.agency/research/ai-prompt-phrasing-study

Frequently asked questions

Does the way you word a question change which brands ChatGPT recommends?

Yes, beyond ordinary variation. Asked the identical question again, ChatGPT named the same first brand 74.5% of the time; with “on a tight budget” added, 21.8%.

Which rewording changes AI recommendations the most?

A budget, among the three we tested. Averaged across ChatGPT, Gemini and Perplexity, it kept the original’s first brand 15.3% of the time and shared 0.326 of its brands with the original, against 68.0% and 0.638 for a rerun.

Does asking an AI for an unbiased answer change its recommendations?

A little. The unbiased request changed the list slightly more than a rerun (an excess of 0.058 in brand overlap), much less than a budget or business context.

Is the budget effect caused by the exact words?

Mostly not. Four ways of asking for something cheap produced similar lists (overlap 0.474 with each other against 0.341 with the original). The exact words mattered more on Perplexity than on ChatGPT or Gemini.

How should a brand track its AI visibility across different phrasings?

Track several wordings of each important buyer question, run each several times, and report visibility per wording and engine, because one brand can appear in every answer to one wording and none to another.

Free strategy call

Want these numbers for your own category?

We run the same measurements for a business’s own buyer questions and competitors. On a free 30-minute call we’ll take a first look and send you a short written read afterward.