{
 "question": "Why do generative AI assistants recommend different brands for the same buyer question, and how much of the difference sits in the evidence they cite, in what they select from it, and in run-to-run variation?",
 "population": "80 buyer questions (10 per industry, 24 naming a place), each asked once of ChatGPT and Gemini (consumer apps via DataForSEO LLM Scraper, US), Perplexity sonar and Claude Haiku 4.5 (APIs with web search) on 2026-09-26: 320 answers. Version 1.1 re-analyzes the same answers; no new engine queries. Stability uses study 8's five runs of 20 of the questions.",
 "inclusion": "All 320 answers; answers that recommend nothing stay in the data and drop out of the pairwise overlap for that pair.",
 "exclusion": "",
 "calculations": [
  "Primary outcome (version 1.1): options each answer recommends, model-coded by Claude Opus from all four answers to a question at once, answers under shuffled letters (engine hidden); entities merged across spelling variants; recommended vs merely mentioned; rank of each recommendation.",
  "Agreement: Jaccard overlap of recommended sets per question and pair; first pick = the answer's rank-1 recommendation. 95% intervals by bootstrap over questions (2,000 resamples, seed 20260926). Joint F-tests for pair (with question fixed effects) and industry.",
  "Evidence ecology: pooled OLS of pairwise recommendation overlap on cited-domain overlap (per 0.1), with question-clustered errors; with controls for place, industry, pair, number of sources and number of recommendations; with question fixed effects (within-question); question random-intercept model; question-level bootstrap of the slope.",
  "Evidence: every cited URL (2,209 unique) fetched over plain HTTP on 2026-09-28; main text extracted; readable = HTTP 200 and at least 400 characters of text that is not a block page. An answer's evidence is usable when at least one cited page and at least half of its cited pages were readable. An option is in the evidence when any of its coded names appears as a whole phrase in the normalized page text.",
  "Decomposition: for each option one assistant recommends and the other does not (both answers with usable evidence): selection = the other's cited pages named it; evidence = only the recommender's cited pages named it; outside = neither's cited pages named it. Sensitivity: only answers whose cited pages were all readable.",
  "Stability: study 8's five runs of 20 questions for ChatGPT, Gemini and Perplexity (rule-based brand mentions): within-assistant vs cross-assistant overlap over all run pairs, typical sets (named in 3 or more of 5 runs), brands named in 3 or more runs by one assistant and never by the other.",
  "Disagreement causes: two model coders (Claude Opus, Claude Sonnet), blind to engine and to each other, coded every question with a fixed nine-code taxonomy from the four answers and their cited domains; agreement as percentage and Cohen's kappa; codes reported where both coders agree.",
  "Version 1.0 rule-based figures (brand mention by exact-name match) are kept in stats.json under statistics."
 ],
 "limitations": [
  "One run per assistant per question for the main sample; stability measured on 20 questions and 3 assistants only.",
  "Cited is not the same as retrieved: assistants may read pages they do not cite, so the evidence measures are lower bounds on what each assistant saw.",
  "Pages were fetched two days after the answers; 35.8% of cited URLs could not be read (blocked, not HTML or too short).",
  "A page naming an option does not mean the page recommends it; name matching can miss variants.",
  "Associations only; no evidence or content was manipulated.",
  "All coding is model-coded (two Claude models); no person coded the sample.",
  "Perplexity and Claude via API; Claude answers from Claude Haiku 4.5."
 ],
 "update": "quarterly",
 "@context": "https://schema.org",
 "@type": "Dataset",
 "name": "Do ChatGPT, Gemini, Perplexity and Claude agree on brands?",
 "version": "1.1",
 "dateCreated": "2026-09-26",
 "creator": {
  "@type": "Organization",
  "name": "Underneath",
  "url": "https://underneath.agency"
 },
 "url": "https://underneath.agency/research/ai-assistants-brand-agreement-study",
 "temporalCoverage": "2026-09-26/2026-09-28",
 "isAccessibleForFree": true,
 "license": "https://creativecommons.org/licenses/by/4.0/",
 "code": "cite/pipeline/s7v11_analyze.py, s7v11_validation.py, s7v11_evidence.py, s7v11_pipeline.py, s7v11_package.py (version 1.1); analyze_agree.py (version 1.0)",
 "variableMeasured": [
  "recommended_options",
  "rec_jaccard",
  "first_pick",
  "cited_domain_jaccard",
  "in_own_cited_pages",
  "selection_evidence_outside",
  "within_vs_cross_run_overlap",
  "disagreement_cause"
 ],
 "dateModified": "2026-09-28",
 "research_questions": {
  "RQ1": "How far do ChatGPT, Gemini, Perplexity and Claude agree on the options they recommend for the same question?",
  "RQ2": "Is the overlap of the sources two assistants cite associated with the overlap of the options they recommend, after accounting for question, place, industry and pair?",
  "RQ3": "Are recommended options named in the pages the assistant cited, and when one assistant recommends an option the other does not, did the other's cited pages name it?",
  "RQ4": "How much of the cross-assistant difference is run-to-run variation? (20 questions, 3 assistants, 5 runs, from study 8.)",
  "RQ5": "What reasons for disagreement can be seen in the answers? (two model coders, fixed taxonomy.)",
  "RQ6": "When the evidence is held constant, how much disagreement remains, and which content changes move which stage? Not answered: needs controlled experiments (planned)."
 },
 "framework": "Recommendation pipeline: question -> retrieval -> cited evidence -> selection -> recommendation (rank) -> justification -> stability. This version observes cited evidence, recommendation, first pick and (on a subset) stability; retrieval beyond citations and justification quality are not observed."
}