{
 "question": "When AI assistants give different answers to the same buyer question, what does a measurement of a brand's AI visibility mean, how is the variation structured (membership, order, first pick, sources, wording), what is it associated with, and how much does one observation get wrong?",
 "population": "20 of the 80 buyer questions from study 7 (at least two per industry, 7 naming a place), each asked five times of ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper, US) and Perplexity sonar (API with web search) on 2026-09-26: 300 answers; the five runs of a question were a median of about 11 minutes apart. Comparison answers: study 18 rewordings (5 shared questions, 3 assistants) and study 17 country answers (10 shared questions, ChatGPT and Gemini, US, GB, CA, AU), collected the same day about 4.4 hours after the runs. Version 1.1 re-analyzes these answers; no new engine queries.",
 "inclusion": "A run is valid when the answer has at least 200 characters. One ChatGPT run (172 characters, an unfinished reply) is excluded from the per-brand frequencies (that question has four valid ChatGPT runs); version 1.0 kept it.",
 "exclusion": "Claude (budget); prompts with fewer than five successful runs.",
 "calculations": [
  "Brands: candidate names from each answer (bold names, list headings, ChatGPT brand entities), classified brand / not brand once per unique string, clustered per question over every answer to that question in studies 8, 17 and 18, then matched by normalized name in the answer text. Rank = order of first mention.",
  "Visibility frequency = runs naming the brand / valid runs, with a Wilson 95% interval. Classes: core 5 of 5, stable 4 of 5, intermittent 2-3, tail 1.",
  "Concentration per question and assistant: entropy H of each brand's share of all mentions across the runs; effective number of brands = exp(H); core and tail sizes.",
  "Run pairs (10 per question and assistant, 6 where a run is invalid): Jaccard overlap of brand sets, top-3 sets, cited registrable domains, cited URLs and answer words; extrapolated rank-biased overlap (p = 0.9); Kendall tau-a among brands both runs name (3 or more shared); same first brand. Volatility = 1 - overlap (order among shared brands: (1 - tau) / 2).",
  "Source-brand coupling: slope of pairwise brand overlap on pairwise domain overlap (per 0.1), pooled with assistant dummies and within question and assistant (fixed effects).",
  "Question characteristics: place-named vs not, industry, list length; share of the variation in mean overlap (20 x 3 grid) due to question, assistant and their interaction (sums of squares).",
  "Time and condition ladder: each later single answer compared with the five runs (mean Jaccard) against the same-batch run-to-run overlap, on the same questions.",
  "Single run vs five: share of the five-run brands one run shows; share of brands missing from a run that were named in 3 or more runs; for brand pairs whose five-run frequencies differ by 0.4 or more, share of runs that name the rarer brand and not the commoner; new brands added by each run, averaged over all run orders.",
  "Intervals: 95% percentile bootstrap resampling the 20 questions (2,000 resamples, seed 20260926), paired across assistants.",
  "Version 1.0 figures are kept in stats.json under version_1_0_statistics."
 ],
 "limitations": [
  "Five runs per question over about 11 minutes: frequencies are coarse (a 3-of-5 brand has a Wilson interval of 23% to 88%) and drift over days and weeks is not measured.",
  "20 questions, 3 assistants; the question characteristics analysis is exploratory.",
  "Outputs only: the data cannot show whether variation comes from retrieval, ranking or generation; cited sources are what the answer shows, not everything retrieved.",
  "The hours-later, rewording and country answers are single answers on 5 or 10 questions; the time and condition steps are small-sample estimates.",
  "Perplexity via API (sonar), not the consumer app; Claude not included.",
  "Brand matching by normalized name can miss variants; rank is order of first mention, not a stated ranking."
 ],
 "update": "Planned: a longitudinal panel of the same questions over several days and weeks (with study 24), 20 runs on a subset to measure how many runs a target precision needs, and a fixed-evidence experiment to separate retrieval from generation.",
 "@context": "https://schema.org",
 "@type": "Dataset",
 "name": "Ask an AI the same question 5 times: do the brands change?",
 "version": "1.1",
 "dateCreated": "2026-09-26",
 "creator": {
  "@type": "Organization",
  "name": "Underneath",
  "url": "https://underneath.agency"
 },
 "url": "https://underneath.agency/research/ai-recommendation-consistency-study",
 "temporalCoverage": "2026-09-26/2026-09-28",
 "isAccessibleForFree": true,
 "license": "https://creativecommons.org/licenses/by/4.0/",
 "code": "cite/pipeline/s8_v11.py, s8_v11_package.py (version 1.1); analyze_agree.py (version 1.0)",
 "variableMeasured": [
  "visibility_frequency",
  "visibility_class",
  "effective_brands",
  "brand_jaccard",
  "rbo",
  "kendall_shared",
  "same_first",
  "top3_jaccard",
  "domain_jaccard",
  "url_jaccard",
  "word_jaccard"
 ],
 "dateModified": "2026-09-28",
 "research_questions": {
  "RQ1": "Existence: how much do the brands recommended by ChatGPT, Gemini and Perplexity change across five identical runs minutes apart?",
  "RQ2": "Structure: how is each brand's visibility distributed across runs (core, stable, intermittent, tail), how concentrated is it, and do membership, order and first pick change in the same way?",
  "RQ3": "Mechanism (observational): is run-to-run change in brands associated with change in cited sources, and does it occur when the cited sources are identical?",
  "RQ4": "Measurement: what does one run miss or get wrong against five, and how precise is a five-run frequency?",
  "RQ5": "Scale: how does same-batch repetition compare with the same question hours later, a reworded question and another country, on the questions shared with studies 17 and 18?",
  "RQ6": "Drift over days and weeks, the number of runs needed for a target precision (measured empirically), and the split between retrieval and generation. Not answered: needs new collection (planned)."
 },
 "framework": "An observed answer is one draw from a distribution. Visibility of a brand for a question on an assistant is estimated as the share of valid runs that name it. Change between runs is split into membership (which brands), order (rank-biased overlap; Kendall tau among shared brands), first pick, top 3, evidence (cited domains and URLs) and wording. Time scales: minutes (this collection), hours (study 17/18 answers about 4.4 hours later), days and weeks (not measured).",
 "datasets": [
  "s7_s8_answers.csv",
  "s8_answers_v11.csv",
  "s8_brand_visibility.csv",
  "s8_run_pairs.csv",
  "s8_question_engine.csv"
 ]
}