---
title: "Does rewording a question change AI brand recommendations?"
description: "In 1,560 AI answers, a tight budget kept the original’s first brand 15.3% of the time, against 68.0% when the same question was simply asked again."
canonical: "https://underneath.agency/research/ai-prompt-phrasing-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI assistants

# Does rewording a question change AI brand recommendations?

Buyers with the same need ask in different words. One says they run a small business, one says money is tight, one asks for a straight answer. This study asks how stable the brands ChatGPT, Gemini and Perplexity recommend are when the question is reworded, and how much of the change comes from the wording rather than from the ordinary variation between two runs of the same question. On 28 September 2026 we asked 20 buyer questions in four wordings, five times each, on all three engines, and three more phrasings of the budget request: 1,560 answers. Every rewording moved the recommendations further than a plain rerun did. A budget moved them most, and an explicit request for an unbiased answer moved them least. That other phrasings change AI recommendations is already established ([arXiv 2605.27440](https://arxiv.org/abs/2605.27440)). What this study adds is a same-batch rerun baseline for every question, a split of the change into its parts, and a test of whether the budget effect comes from the constraint or from the particular words.

## The short version

1. Asking the identical question again kept the same first brand 68.0% of the time. Adding “on a tight budget” kept it 15.3%, adding “I run a small business with about 10 employees” 40.6%, and asking for “an honest, unbiased answer” 54.9%.
2. Every rewording changed the brand list by more than a rerun. Brand overlap with the original (Jaccard) was 0.638 for a rerun, 0.547 for the unbiased request, 0.389 for business context and 0.326 for a budget. The excess over rerun variation was 0.273 for a budget (95% interval 0.220 to 0.322) and 0.058 for the unbiased request (0.036 to 0.080).
3. The wording explained 46.7% of the variation in brand lists between the budget and original answers to the same question, and 21.6% for the unbiased request; the rest is run-to-run variation. By chance alone it would explain about 0.11.
4. The budget effect comes mainly from the constraint, not the exact words. Four different ways of saying “cheap” overlapped with each other at 0.474, far closer to a same-phrasing rerun (0.563) than to the original question (0.341). The words mattered more on Perplexity than on ChatGPT or Gemini.
5. The top brand is not protected. A brand named first in the original survived a rerun 94.3% of the time but a budget rewording only 54.5%. Brands that came in on a budget were given a price or value reason 55.9% of the time (model-coded).

## Research questions

| Question | Answered here? |
|---|---|
| RQ1. Does adding context, a constraint or a meta-instruction change the brands recommended? | Yes |
| RQ2. Is that change larger than rerunning the identical question? | Yes, for every question, engine and wording |
| RQ3. How do the lists change: membership, first brand, top three, order, length? | Yes |
| RQ4. Does sensitivity differ by engine and category? | Yes, with a mixed model |
| RQ5. Which brands hold their place and which appear only under one wording? | Yes |
| RQ6. Does the budget effect come from the constraint or the words? | Yes, with four phrasings |

## Hypotheses and results

| Hypothesis | Result |
|---|---|
| H1. Adding decision information lowers overlap with the original | Supported |
| H2. A budget changes the list more than an unbiased-answer request | Supported |
| H3. Rewordings change the list more than a rerun | Supported for all three |
| H4. Sensitivity differs by engine | Partly |
| H5. Some brands are far more wording-sensitive than others | Supported |
| H6. Rewording removes lower-ranked brands more than top ones | Not for a budget |

H1: business context (0.389) and budget (0.326) against a rerun (0.638). H2: the intervals for budget (0.272 to 0.386) and the unbiased request (0.503 to 0.590) do not overlap. H3: the excess over reruns is above zero for all three rewordings, though small for the unbiased request. H4: the raw amount of change did not differ by engine once wording was accounted for (p = 0.304), but measured against each engine’s own rerun variation Perplexity was the most wording-sensitive (p = 0.006). H5 and H6 are covered under the brand findings below.

## A model of what moves a recommendation

We treat an AI recommendation as the result of several inputs: the buyer’s underlying need (the question), the wording, the engine, the date and the run. This study changes the wording while holding the question, engine and date fixed, and measures run-to-run variation directly by asking each wording five times. A brand therefore does not have one AI visibility score. It has a visibility for a question, in a wording, on an engine, at a time.

The three rewordings are different kinds of change, not three paraphrases.

| Rewording | Kind of change | Example |
|---|---|---|
| Business context | Who is asking | “I run a small business with about 10 employees. What is the best air fryer?” |
| Budget | A constraint | “What is the best air fryer on a tight budget?” |
| Unbiased answer | How to answer | “I’m comparing options and want an honest, unbiased answer, not a sales pitch: what is the best air fryer?” |

## What we measured

We took the 20 national buyer questions used in the first version of this study, drawn from our [four-assistant study](https://underneath.agency/research/ai-assistants-brand-agreement-study), such as “What is the best air fryer?” and “Which email marketing platform is best for ecommerce?”.

- **Main design (1,200 answers):** each question in the original wording and the three rewordings, five runs each, on ChatGPT and Gemini (their consumer apps) and Perplexity (sonar with web search), United States, 28 September 2026.
- **Budget phrasings (360 answers):** the budget request stated three other ways, two runs each on each engine: “What is the best affordable air fryer?”, “What is the best air fryer for someone trying to spend less?” and “Keeping the total price as low as possible is my priority. What is the best air fryer?”.
- **Same batch:** the answers to all wordings of a question on one engine were collected side by side, a median of 0.1 minutes apart. The whole collection took 0.4 hours, so time cannot explain the differences between wordings.

Brands were identified as in the four-assistant study: names in bold or headings, classified as brand or not, name variants merged per question, then matched in each answer in order of first mention. For each pair of answers we measured brand overlap (Jaccard: brands in both divided by brands in either), whether the first brand matched, how many of the first three were shared, rank-biased overlap (an agreement score that weights the top of the list most), the share of the original’s brands kept, the share of new brands, the order of shared brands (Kendall’s tau) and the change in list length. The rerun baseline compares the five runs of the same wording with each other. Every 95% interval comes from resampling the 20 questions, because answers to one question are not independent.

## Findings

### Every rewording moved the list further than a rerun

| Against the original | Overlap | First brand | Excess | Interval |
|---|---|---|---|---|
| Same wording, run again | 0.638 | 68.0% | | |
| Unbiased answer | 0.547 | 54.9% | 0.058 | 0.036 to 0.080 |
| Business context | 0.389 | 40.6% | 0.194 | 0.154 to 0.240 |
| Budget | 0.326 | 15.3% | 0.273 | 0.220 to 0.322 |

Overlap is brand overlap (Jaccard), first brand the share with the same first brand, and interval the 95% interval of the excess. Excess change is the drop in brand overlap beyond what reruns of the two wordings show on their own. The excess was positive in 93.3% of the 60 question-and-engine pairs for a budget, 91.7% for business context and 76.7% for the unbiased request. For the first brand, the budget excess was 48.8 percentage points.

### How much of the variation is the wording

Within each question and engine we split the variation in brand lists into a part explained by the wording and a part left over from run to run (a permutation analysis of variance on brand-list distances).

| Wording | Explained | Significant |
|---|---|---|
| Unbiased answer | 21.6% | 28.3% |
| Business context | 38.6% | 68.3% |
| Budget | 46.7% | 81.7% |

Explained is the share of variation explained by the wording; significant is the share of the 60 question-and-engine pairs where that effect is significant at the 5% level. With two wordings and ten answers, chance alone would give about 0.11. Taking all four wordings together, the wording explained 50.4% of the variation (95% interval 46.6% to 53.7%) and was significant at the 5% level in 95.0% of question-and-engine pairs.

A mixed model of the change per question, engine and rewording, with a random effect for each question, confirms the ordering: relative to the unbiased request, a budget raised the set-change score by 0.221 (0.164 to 0.277) and business context by 0.158 (0.102 to 0.214), and the wording as a whole was highly significant (p < 0.001). Differences between questions accounted for about a quarter of the remaining variance (0.255).

### Engines differ in noise more than in sensitivity

| Brand overlap with the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| ChatGPT | 0.642 | 0.555 | 0.408 | 0.363 |
| Gemini | 0.545 | 0.469 | 0.327 | 0.305 |
| Perplexity | 0.727 | 0.616 | 0.431 | 0.310 |

| Same first brand as the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| ChatGPT | 74.5% | 57.4% | 48.5% | 21.8% |
| Gemini | 50.5% | 40.4% | 33.3% | 13.6% |
| Perplexity | 79.0% | 66.8% | 40.0% | 10.4% |

Gemini’s lists change the most in absolute terms, but much of that is its own run-to-run variation: two runs of the identical question shared only 0.545 of their brands. Perplexity is the most repeatable (0.727) and so the most wording-sensitive once its low rerun variation is taken into account: the budget excess was 0.414 on Perplexity against 0.226 on ChatGPT and 0.179 on Gemini, and in the model of excess change Perplexity added 0.113 (0.069 to 0.157). The raw amount of change did not depend on the engine and wording together (p = 0.304).

### How the lists change

| Measure (all engines) | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| Brand overlap | 0.638 | 0.547 | 0.389 | 0.326 |
| Top three shared | 75.5% | 66.7% | 51.6% | 33.1% |
| Rank-biased overlap | 0.794 | 0.713 | 0.554 | 0.418 |
| Original brands kept | | 72.0% | 49.2% | 43.7% |
| New brands | | 31.0% | 39.3% | 48.2% |
| Order of shared brands | | 0.495 | 0.361 | 0.100 |
| Change in list length | | +0.24 | −1.00 | −1.12 |

Rank-biased overlap and order are scores from 0 to 1, where 1 is identical. A budget replaced about half the list and scrambled the order of the brands that stayed: the order score of 0.100 (interval −0.03 to 0.235) is not distinguishable from random. Business context and a budget both shortened the list by about one brand; the unbiased request slightly lengthened it.

### The top brand is not protected

For each brand in an original answer we asked how often it appeared in a rerun, and how often in a reworded answer.

| Place in the original | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|
| First | 94.3% | 87.6% | 69.1% | 54.5% |
| Second or third | 85.0% | 80.3% | 58.4% | 42.8% |
| Fourth or lower | 65.4% | 58.5% | 35.7% | 34.2% |

Based on 293 first-placed, 574 second- or third-placed and 903 lower-placed brand mentions. Lower-placed brands are less stable in every condition, but a budget costs the first brand about as much as the rest of the list (it keeps 54.5% of its 94.3%). We expected rewording to trim mainly the tail of the list (H6); for a budget it does not. Business context does hit the tail harder.

### Which brands hold their place

For each brand, question and engine we counted how often the brand appeared in the five runs of each wording (1,074 brand-question-engine combinations).

| Brand pattern | Share |
|---|---|
| Named in only one or two of about 20 answers | 40.6% |
| Leaves when a budget is added | 15.0% |
| No clear pattern | 14.5% |
| Enters when a budget is added | 9.4% |
| Stable: in at least 60% of runs of every wording | 7.4% |
| Enters with business context | 5.0% |
| Leaves with business context | 4.9% |
| Moves with the unbiased request | 3.2% |

Of the 337 brands named in at least 60% of the original’s runs, 78.6% held that level under the unbiased request, 51.3% under business context and 42.7% under a budget, and 47.8% were named in at most one of the five budget runs. Among the 638 brands named in three or more answers, 88.6% had a visibility that differed by 40 points or more between their best and worst wording. Engine-specific brands, named in half of one engine’s answers and almost never by the other two, were 10.5% of the 143 brands that reached that level anywhere.

### A brand does not have one AI visibility score

Perplexity, “Which email marketing platform is best for ecommerce?”, runs out of five that named each brand.

| Brand | Original | Context | Budget | Unbiased |
|---|---|---|---|---|
| Omnisend | 5 | 5 | 5 | 5 |
| Klaviyo | 5 | 5 | 0 | 5 |
| Mailchimp | 5 | 5 | 0 | 2 |
| Drip | 5 | 0 | 0 | 2 |
| ActiveCampaign | 4 | 0 | 0 | 5 |
| MailerLite | 0 | 3 | 5 | 4 |

Klaviyo was named in every Perplexity answer except the budget ones. On the same budget question ChatGPT named Klaviyo in all five runs and Gemini in one. Asked for the best air fryer, Perplexity named Cosori, TurboBlaze, Ninja and Instant Vortex Plus in all five original runs; on a tight budget it led with Gourmia and the Chefman TurboFry 2 Quart, and named Ninja in none. A single “AI visibility” figure for any of these brands would average over wordings that give 0 and 5 out of five.

### The constraint matters more than the words

To test whether the budget effect comes from the constraint or from the phrase “tight budget”, we asked the same constraint four ways.

| Comparison (all engines) | Brand overlap | Same first brand |
|---|---|---|
| Original wording, run again | 0.638 | 68.0% |
| Same budget phrasing, run again | 0.563 | 60.1% |
| Two different budget phrasings | 0.474 | 48.6% |
| A budget phrasing and the original | 0.341 | |

The phrasings agreed with each other much more than any of them agreed with the original: overlap with the original was 0.326 for “tight budget”, 0.370 for “affordable”, 0.353 for “someone trying to spend less” and 0.338 for “lowest total price is my priority”. Brands that came in with “tight budget” also appeared with each other phrasing 69.6% of the time. The words still mattered a little. The gap between a same-phrasing rerun and a different phrasing was 0.090 across engines (0.061 to 0.117), but it was small on ChatGPT (0.031) and Gemini (0.042) and larger on Perplexity (0.196), whose answers depend more on the exact search terms. A recent study of API models found that the prompt string, more than the buyer’s intent, decided which brands surfaced. In these consumer assistants, the intent (a price limit) did most of the work.

### What changes with the wording: sources and prices

Each engine cites web pages with its answers, so we can see whether a reworded question also draws on different sources.

| Domains | Cites | Rerun | Unbiased | Context | Budget |
|---|---|---|---|---|---|
| ChatGPT | 97.5% | 0.632 | 0.390 | 0.292 | 0.282 |
| Gemini | 95.2% | 0.325 | 0.190 | 0.107 | 0.147 |
| Perplexity | 100.0% | 0.986 | 0.820 | 0.517 | 0.595 |

Overlap of cited domains with the original; “cites” is the share of answers citing any source. Rewording changed the cited sources far more than rerunning did, and the questions whose sources changed most were the ones whose brands changed most (Spearman correlation 0.35 on ChatGPT, 0.344 on Gemini and 0.427 on Perplexity, each p < 0.01). Perplexity cited almost the same sites on every rerun (0.986), so its brand changes follow the changed search. The answers also changed in content: 84.7% of budget answers quoted a dollar figure, against 33.0% of original answers, 37.3% with business context and 43.0% with the unbiased request.

### Why brands came in with a budget

For the 10 questions where a budget changed the list most, we took each brand that appeared in the first budget answer but not in the first original answer (102 brands) and coded the main reason the answer itself gave for it. Two independent Claude model coders agreed on 87.3% (Cohen’s kappa 0.799); a third model pass settled the 13 disagreements. These are model-coded, not human-validated.

| Reason given for a new brand | Share |
|---|---|
| Price or value | 55.9% |
| Features | 21.6% |
| Suited to a type of buyer | 7.8% |
| Quality or reliability | 6.9% |
| Named as an alternative, no reason | 5.9% |
| No reason | 2.0% |

Price or value was the main reason for 71.0% of new brands on Perplexity, 55.9% on ChatGPT and 43.2% on Gemini. The change was mostly explained, not arbitrary: the assistants brought in brands they could justify on price.

### Where rewording matters most

| Category | Questions | Rerun | Context | Budget |
|---|---|---|---|---|
| Consumer products | 5 | 0.555 | 0.386 | 0.194 |
| Business software | 5 | 0.708 | 0.526 | 0.363 |
| Insurance and finance | 7 | 0.610 | 0.289 | 0.341 |
| Health and legal services | 3 | 0.725 | 0.398 | 0.451 |

Brand overlap with the original. Consumer products were the most budget-sensitive: robot vacuums (a set-change score of 0.844, where 1 means no brand in common), noise-cancelling headphones (0.828) and air fryers (0.827). The least sensitive were online LLC formation (0.355), car insurance for young drivers (0.486), whose original question already asks for the cheapest, and robo-advisors (0.502). Questions with more stable reruns were somewhat less budget-sensitive (Spearman 0.29, p = 0.025). The category groups are small, so these differences are descriptive.

### Two days apart, and the first version

Answers to the original wording on 26 September (from the first version of this study) overlapped with the 28 September runs about as much as two same-day runs did: 0.642 against 0.642 on ChatGPT, 0.546 against 0.545 on Gemini, and 0.688 against 0.727 on Perplexity. Over two days, time added little beyond run-to-run variation.

| Average of three engines | v1.0 | v1.1 |
|---|---|---|
| Same first brand, budget | 12.0% | 15.3% |
| Same first brand, business context | 40.0% | 40.6% |
| Same first brand, unbiased | 55.0% | 54.9% |
| Brand overlap, budget | 0.366 | 0.326 |
| Brand overlap, business context | 0.412 | 0.389 |
| Brand overlap, unbiased | 0.574 | 0.547 |

Version 1.0 used one run per wording on 26 September; version 1.1 five runs per wording on 28 September. The first version’s ordering holds on all three engines. One difference: in version 1.0, ChatGPT’s budget answers overlapped with the original more than its business-context answers did (0.454 against 0.402). With five runs, the budget is the largest change on every engine.

## A prompt sensitivity index

Because no single number captures how a list changes, we report sensitivity as three scores, each from 0 (no change) to 1 (complete change), between the original and the reworded answers.

| Wording | Set change | First-brand change | Rank change |
|---|---|---|---|
| Same wording, run again | 0.362 | 0.32 | 0.206 |
| Unbiased answer | 0.453 | 0.451 | 0.287 |
| Business context | 0.611 | 0.594 | 0.446 |
| Budget | 0.674 | 0.847 | 0.582 |

Set change is 1 minus brand overlap, first-brand change is 1 minus the same-first-brand rate, and rank change is 1 minus rank-biased overlap. The first row is the floor that ordinary variation sets. A tracker that reports one prompt run once cannot tell a wording effect from that floor.

## How this compares with other studies

| Study | What changed | Engines | Runs | Main figure |
|---|---|---|---|---|
| Aggarwal et al. 2024 (GEO) | The source pages | Generative engines | Controlled | Page edits change how visible a source is |
| Chen et al. 2025 | Engine, language, vertical, paraphrase | Several AI search engines | Controlled | Cited sources differ by engine and phrasing |
| Jack et al. 2026, persona | A persona prefix | OpenAI, Anthropic APIs | 2,000 | Jaccard lowered by 0.12 to 0.20 |
| Jack et al. 2026, paraphrase | Cosmetic and constraint paraphrases | OpenAI, Anthropic APIs | About 12,000 | 0.288 and 0.135, against 0.50 to 0.61 for reruns |
| Rankshift 2026 | Seven near-synonym CRM prompts | Several | 1,176 each | HubSpot at 98% on every variant |
| Tannenbaum 2026 | None (monitoring data) | GPT, Gemini | 34,960 | Mentions depend on the brand being retrieved |
| This study | Context, constraint, meta-instruction, four budget phrasings | ChatGPT, Gemini, Perplexity apps | 5 per wording | 0.326 for a budget against 0.638 for reruns |

The API study of paraphrases found constraint-adding rewordings overlapping at 0.135 against 0.50 to 0.61 for reruns. Our budget rewording (0.326 against 0.638) points the same way, less extremely. The clearest difference is on cosmetic paraphrase: there, different wordings of the same intent overlapped at 0.288; here, different wordings of the same budget constraint overlapped at 0.474, close to a rerun on ChatGPT and Gemini. Persona conditioning lowered overlap by 0.12 to 0.20, and our business-context wording, a short persona, lowered it by more (0.194 beyond reruns). A study of monitoring data found that brands are mentioned mainly when their own pages appear in what the engine retrieves. Our finding that brand changes follow source changes is consistent with that.

Sources: [arXiv 2311.09735](https://arxiv.org/abs/2311.09735); [arXiv 2509.08919](https://arxiv.org/abs/2509.08919); [arXiv 2605.30207](https://arxiv.org/abs/2605.30207); [arXiv 2605.27440](https://arxiv.org/abs/2605.27440); [Rankshift](https://www.rankshift.ai/blog/impact-of-prompt-phrasing-on-ai-brand-visibility/); [arXiv 2609.23162](https://arxiv.org/abs/2609.23162).

## Observed, inferred and unknown

**What we observe.** Asked the same question in the same batch, the three assistants return lists that vary from run to run, and adding context or a constraint changes the list well beyond that. A budget changes the first brand in most answers, replaces about half the brands and scrambles the order of the rest. Several ways of stating a budget lead to similar lists. The cited sources change along with the brands, and the new brands come with price reasons.

**What we infer.** The assistants treat a budget or a stated situation as a different request and search for it differently, rather than reacting to particular words. AI brand visibility is conditional on how the need is expressed, so it should be measured across wordings and repeated runs, not from one prompt run once.

**What remains unknown.** Whether the new recommendations are better or worse for the buyer: we measured difference, not quality. How much answers drift over weeks rather than two days. Whether logged-in history or personalization adds more variation. Which brand characteristics (price level, size, review volume) predict holding a place across wordings: the brand-level data is published for that analysis, but we have not collected those attributes.

## What this means

The points below are interpretation. They follow from the findings but were not tested.

- **Report visibility for a question in a wording, not one score.** A brand that appears in every answer to “What is the best email marketing platform for ecommerce?” can be absent from every answer that adds a budget. Measure each important wording separately.
- **Run each prompt several times before reading a change.** Two runs of the identical question shared 0.638 of their brands. A difference smaller than that is often noise.
- **Match the buyer’s constraint, not their exact phrase.** Different ways of asking for something cheap led to similar lists, so content that states prices and value plainly serves all of them.
- **Being first in the default answer does not carry over.** The first-named brand kept its place in only about half of budget answers.

## Where this sits in GEO research

Generative engine optimization began by asking whether changing a source page changes how visible it is in AI answers. Later work showed that AI search engines differ by language, vertical and phrasing. A 2026 survey of the field names repeated runs, paraphrase, denominators and causal measurement as unsolved measurement problems. This study takes up two of them: it separates wording effects from rerun variation with a same-batch baseline, and it shows that a brand’s visibility has no meaning without saying for which wording. It does not show how to change that visibility. The next collection repeats the main design on a new date to separate time from run-to-run variation.

## Methodology

- **Questions:** 20 national buyer questions from the four-assistant study (every second national question), each in the original wording and three rewordings; three further budget phrasings.
- **Engines:** ChatGPT and Gemini consumer apps via DataForSEO LLM Scraper; Perplexity (sonar) via DataForSEO LLM Responses API with web search; United States; English.
- **Collection:** 1,560 requests on 28 September 2026 between 08:57 and 09:19 UTC: 1,200 in the main design (five runs per wording) and 360 for the budget phrasings (two runs each). All 1,560 returned; six unfinished ChatGPT answers under 200 characters are kept as outcomes but excluded from brand measures, and 18 answers named no brand. Version 1.0 answers (26 September 2026, one run per wording) are used for the cross-date and version comparisons.
- **Brands:** candidate names from bold text and headings (and ChatGPT’s brand panels); each unique string classified as brand or not (study 7 decisions reused; new strings first filtered by rule for prices, pronouns and long phrases, then classified by Claude model coders, with directive fragments and URLs corrected by rule); name variants merged per question; matched in each answer in order of first mention. Names that are also ordinary words (such as Choice, Fidelity or Travelers) count only where the answer lists them as an item, and review sites cited as sources are not counted as brands.
- **Measures:** Jaccard similarity, same first brand, top-three overlap, extrapolated rank-biased overlap (persistence 0.9), retention, replacement, Kendall’s tau on shared brands (three or more), list length. Rerun baseline from all pairs of the five runs; between-wording figures from all 25 pairs of original and reworded runs.
- **Variance and models:** permutation analysis of variance on Jaccard distance per question and engine (499 permutations); linear mixed model with wording and engine as fixed effects and question as a random intercept; likelihood-ratio tests.
- **Reasons for change:** 102 brands entering in the first budget answer against the first original answer, for the 10 most budget-sensitive questions; two independent model coders, third-pass adjudication.
- **Uncertainty:** 95% bootstrap intervals resampling the 20 questions (2,000 resamples, seed 20260928).
- **Update schedule:** quarterly; next collection repeats the main design on a new date.

## Limitations

- Three engines, 20 questions, the United States and English, with the main design on one day. Other categories, countries and dates may differ.
- ChatGPT and Gemini answers come from their consumer apps through a data provider, without a signed-in history; Perplexity is the API with web search. Personalization is not measured.
- Four wordings and four budget phrasings are a small sample of how buyers ask. The business-context wording was added to every question, including consumer ones where it is less natural.
- Five runs per wording estimate run-to-run variation with some noise, and two runs per budget phrasing is thin.
- Brand identification is automatic and can merge or split names; the reasons for change are model-coded, not human-validated.
- We change the question, not the model or its retrieval, so the source analysis shows association, not cause.
- Different recommendations are not necessarily better or worse ones.

## What changed in version 1.1

Version 1.0 (26 September 2026) asked each rewording once and compared it with the original from our four-assistant study, with a rerun comparison on only 5 questions from a batch collected hours earlier. Following an external review, version 1.1 (28 September 2026) is a new collection of 1,560 answers. It adds five same-batch runs of every wording for all 20 questions, measures of how lists change (not only Jaccard), a variance split and mixed model with intervals clustered by question, brand-level patterns, model-coded reasons for change, a test of four budget phrasings, hypotheses taken from the review before the new data were collected, and a comparison with the literature. Brand matching now ignores ordinary words and review sites that the earlier method counted as brands. The version 1.0 figures are kept in the data files and compared above; its conclusions hold.

## Data and downloads

- Every answer with its brands in order, first brand and cited domains: [s18v11_answers.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_answers.csv) and [JSON](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_answers.json)
- Every question, engine and wording comparison with all measures: [s18v11_cells.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_cells.csv)
- Each brand’s appearance rate by wording: [s18v11_brand_fate.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_brand_fate.csv)
- The budget-phrasing comparison: [s18v11_paraphrase_cells.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_paraphrase_cells.csv)
- Reason codes for brands entering with a budget: [s18v11_reason_codes.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18v11_reason_codes.csv)
- Every statistic on this page: [stats.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/stats.json); version 1.0: [stats_v1.0.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/stats_v1.0.json) and [s18_phrasing_answers.csv](https://underneath.agency/research-data/ai-prompt-phrasing-study/s18_phrasing_answers.csv)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-prompt-phrasing-study/methodology.json)
- Charts: [rewording against rerun](https://underneath.agency/research-data/ai-prompt-phrasing-study/rewording-vs-repeat-overlap.svg); [share of variation from wording](https://underneath.agency/research-data/ai-prompt-phrasing-study/share-of-variation-from-wording.svg)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *Does rewording a question change AI brand recommendations?* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-prompt-phrasing-study

## Frequently asked questions

### Does the way you word a question change which brands ChatGPT recommends?

Yes, beyond ordinary variation. Asked the identical question again, ChatGPT named the same first brand 74.5% of the time; with “on a tight budget” added, 21.8%.

### Which rewording changes AI recommendations the most?

A budget, among the three we tested. Averaged across ChatGPT, Gemini and Perplexity, it kept the original’s first brand 15.3% of the time and shared 0.326 of its brands with the original, against 68.0% and 0.638 for a rerun.

### Does asking an AI for an unbiased answer change its recommendations?

A little. The unbiased request changed the list slightly more than a rerun (an excess of 0.058 in brand overlap), much less than a budget or business context.

### Is the budget effect caused by the exact words?

Mostly not. Four ways of asking for something cheap produced similar lists (overlap 0.474 with each other against 0.341 with the original). The exact words mattered more on Perplexity than on ChatGPT or Gemini.

### How should a brand track its AI visibility across different phrasings?

Track several wordings of each important buyer question, run each several times, and report visibility per wording and engine, because one brand can appear in every answer to one wording and none to another.

## Related research

- [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)
- [Same question, four countries: do AI recommendations change?](https://underneath.agency/research/ai-recommendations-by-country-study)
- [The hidden searches AI assistants run before they answer](https://underneath.agency/research/ai-hidden-searches-study)
- [Do ChatGPT, Gemini, Perplexity and Claude agree on brands?](https://underneath.agency/research/ai-assistants-brand-agreement-study)

## Related guides

- [Do AI search engines just agree with however a question is phrased?](https://underneath.agency/resources/do-ai-search-engines-agree-with-leading-questions)
- [How should you design prompts and runs to track AI visibility?](https://underneath.agency/resources/how-to-design-ai-visibility-tracking)
- [Can I trust a one-time AI visibility report for my brand?](https://underneath.agency/resources/one-time-ai-visibility-report)
- [How many prompts do we need to track to measure AI visibility?](https://underneath.agency/resources/how-many-prompts-to-track-ai-visibility)
- [Does the way customers phrase a question change which sources AI search cites?](https://underneath.agency/resources/does-question-phrasing-change-ai-sources)

---

This is the Markdown twin of https://underneath.agency/research/ai-prompt-phrasing-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
