{
 "question": "How stable are the brands ChatGPT, Gemini and Perplexity recommend when a buyer rewords the same question, and how much of the change is due to the wording rather than ordinary run-to-run variation?",
 "population": "20 national buyer questions from study 7. Block A: each question in the original wording and three rewordings, five runs each, of ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper) and Perplexity (sonar with web search via DataForSEO LLM Responses), United States, 28 September 2026 (1,200 requests). Block B: three further phrasings of the budget constraint, two runs each (360 requests). Version 1.0 data (26 September 2026, one run per wording) kept as the first study and for the cross-date comparison.",
 "inclusion": "All answers of at least 200 characters.",
 "exclusion": "Unfinished answers under 200 characters (listed in stats.json) are excluded from brand measures and counted as outcomes.",
 "calculations": [
  "Brands: candidate extraction (bold and heading text, ChatGPT brand entities), classification of each unique string (study 7 decisions reused; new strings rule-filtered for prices, pronouns and long phrases, then classified by Claude model coders), clustering of name variants per question, exact matching of normalized names; order = first mention.",
  "Pair measures between two answers: Jaccard similarity of brand sets; same first brand; top-three overlap (share of the reference answer's first three brands in the other answer's first three); extrapolated rank-biased overlap (p = 0.9); retention (share of reference brands kept); replacement (share of the other answer's brands that are new); Kendall's tau on shared brands (three or more shared); change in list length.",
  "Repeat baseline: all pairs among the five runs of one wording. Between-wording: all 25 pairs of original and reworded runs. Rerun ceiling = mean of the two within-wording Jaccards; excess change = rerun ceiling minus between-wording Jaccard.",
  "Prompt sensitivity index (vector): PSI_set = 1 - Jaccard, PSI_top1 = 1 - same-first-brand rate, PSI_rank = 1 - rank-biased overlap, each between original and reworded runs.",
  "Variance decomposition: PERMANOVA on Jaccard distance within each question and engine; R-squared = share of brand-set variation explained by wording; 499 label permutations for p; chance level = (groups - 1)/(answers - 1).",
  "Mixed model: PSI_set and excess change per question x engine x rewording with wording and engine as fixed effects and a random intercept per question (statsmodels MixedLM, ML); likelihood-ratio tests for wording and for wording x engine.",
  "Brand fate: appearance rate of each brand in the five runs of each wording, per question and engine; stable = at least 60% in all four wordings; enters/leaves = at least 60% under one wording and at most 20% under the original (or the reverse); volatile = named in at most 2 of about 20 answers.",
  "Reason for change: entering brands in run 1 budget versus run 1 original answers for the 10 questions with the largest budget change, coded by two independent Claude model coders into seven reasons, disagreements settled by a third model pass; agreement and Cohen's kappa reported. Model-coded, not human-validated.",
  "Uncertainty: 95% bootstrap intervals resampling the 20 questions (2,000 resamples, seed 20260928)."
 ],
 "limitations": [
  "Three engines, 20 questions, United States, English, one collection day for the factorial; results may not hold for other categories, countries or dates.",
  "ChatGPT and Gemini are consumer-app answers collected by a data provider without a logged-in history; Perplexity is the API with web search. Personalization is not measured.",
  "Four wordings plus three budget phrasings are a small sample of how buyers ask.",
  "Five runs per cell estimate run-to-run variation with some noise; two runs per budget phrasing is thin.",
  "Brand identification is automatic and can merge or split names; reason codes are model-coded.",
  "Observational with respect to the engines: we manipulate the question, not the retrieval or model, and cannot see why a system chose a brand.",
  "Different recommendations are not better or worse recommendations: quality was not assessed."
 ],
 "update": "Quarterly; next collection repeats Block A on a new date to separate time from run-to-run variation.",
 "@context": "https://schema.org",
 "@type": "Dataset",
 "name": "Does rewording a question change AI brand recommendations?",
 "version": "1.1",
 "dateCreated": "2026-09-26",
 "creator": {
  "@type": "Organization",
  "name": "Underneath",
  "url": "https://underneath.agency"
 },
 "url": "https://underneath.agency/research/ai-prompt-phrasing-study",
 "temporalCoverage": "2026-09-26/2026-09-28",
 "isAccessibleForFree": true,
 "license": "https://creativecommons.org/licenses/by/4.0/",
 "code": "cite/pipeline/ (collection, validation and analysis scripts)",
 "variableMeasured": [
  "collected",
  "version_1_0_collected",
  "reanalyzed",
  "bootstrap",
  "answers",
  "questions",
  "timing",
  "brands_per_answer_median",
  "brands_per_answer_mean",
  "by_engine",
  "r2_chance_level",
  "mixed_model",
  "rank_layer_survival_pct",
  "rank_layer_n",
  "brand_fate",
  "conditional_visibility_examples",
  "by_group",
  "baseline_stability_vs_budget_overlap_spearman",
  "budget_change_by_question",
  "sources",
  "price_language_pct",
  "dollar_figures_pct",
  "paraphrase_family",
  "cross_date",
  "v1_0_reextracted_jaccard",
  "hypotheses",
  "reason_codes",
  "percent_view",
  "psi_rerun_baseline"
 ],
 "dateModified": "2026-09-28",
 "research_questions": {
  "RQ1": "Does adding context, a constraint or a meta-instruction to a buyer question change the brands recommended?",
  "RQ2": "Is that change larger than the change between two runs of the identical question?",
  "RQ3": "How do the recommendations change: set membership, first brand, top three, order, list length?",
  "RQ4": "Does sensitivity differ by engine and by category?",
  "RQ5": "Which brands hold their place across wordings and which appear only under one wording?",
  "RQ6": "Does the budget effect come from the constraint or from the particular words used to state it?"
 },
 "hypotheses": {
  "H1": "Rewordings that add decision information lower brand-set overlap with the original",
  "H2": "A budget constraint changes the set more than a request for an unbiased answer",
  "H3": "Substantive rewordings change the set more than same-wording reruns",
  "H4": "Sensitivity differs by engine",
  "H5": "Some brands are far more wording-sensitive than others",
  "H6": "Rewording removes lower-ranked brands more than top-ranked ones"
 },
 "framework": "Recommendation R = f(question/intent, wording, engine, time, run). Wording is manipulated; run-to-run variation is estimated from five identical runs per cell; time from the 26 September (v1.0) versus 28 September collections. A brand's AI visibility is treated as conditional on wording and engine, not as one score.",
 "prompt_taxonomy": {
  "context": "adds who is asking (I run a small business with about 10 employees)",
  "budget": "adds a decision constraint (on a tight budget)",
  "conversational": "adds a meta-instruction about the answer style (honest, unbiased, not a sales pitch)"
 }
}