{
 "question": "Is a brand's public entity representation (Wikipedia, Wikidata, Organization schema, sameAs) associated with being recommended by several AI assistants for the same buyer question, and does the association remain after accounting for brand prominence, independent coverage, question, place and assistant?",
 "population": "Every option named in 320 answers (ChatGPT and Gemini consumer apps via DataForSEO, Perplexity sonar and Claude Haiku 4.5 APIs with web search; 80 US buyer questions; 26 September 2026): 1,419 entity-question pairs, model-coded by Claude Opus in study 7 v1.1 (mentioned vs recommended, rank), 1,282 distinct brand keys. Version 2.0 re-analyzes these answers; no new AI queries.",
 "inclusion": "Unit of analysis: entity x question x assistant (5,676 rows). Version 1.0's 914 extracted brands supply 966 of the pairs; the other 453 pairs are options the automatic extractor missed, given the same signals with the same rules (s12v2_extend.py).",
 "exclusion": "",
 "calculations": [
  "Outcomes: mention (named at all), recommendation (model-coded as recommended), top-3 rank among recommendations, number of assistants recommending the entity for the question (0-4), own website cited, named in the assistant's own cited pages, and share of five runs naming the brand (study 8).",
  "English Wikipedia article (primary, v2): the Wikidata match has an English sitelink (v1 method) OR the entity's own name resolves to a Wikipedia article whose short description fits the industry (redirects to a different, parent title count only at parent level). Article-or-parent adds articles for the parent brand the coder recorded or a redirect target.",
  "Wikidata: strict v1 rule; 119 new matches reviewed one by one, 13 rejected. Title matches: all 40 additions reviewed; 6 titles rejected by review.",
  "Prominence controls: Tranco top-1M rank of the brand's website (list 64X3X, 6 - log10 rank; 0 and a flag when not listed) and log10(1 + English Wikipedia articles that mention the name). News coverage (GDELT) was attempted but refused.",
  "Independent coverage: number of distinct domains, other than the brand's own site, among the 1,419 readable pages cited anywhere in study 7 that name the brand (log10(1+n)). Inside the assistants' own evidence, so partly an outcome of the same process.",
  "Models: logistic regression with question fixed effects and assistant dummies, standard errors clustered by brand, fitted as a ladder: signal only (M0), plus prominence (M1), plus independent coverage (M2). Average marginal effects by recomputation. Robustness: Bayesian mixed logit with random intercepts for brand and question (variational Bayes).",
  "Interactions: likelihood-ratio tests for signal x assistant, signal x national/local, signal x question wording (plain, constrained, attribute; a rule on the wording).",
  "Tiers (homepage-checked national brands): weak, partial, strong (Wikidata + Organization schema + sameAs), strong and corroborated (strong + above-median independent coverage).",
  "Matched pairs: within the same national question, each entity with an article paired with the nearest entity without one on the three prominence and coverage measures (standardized distance <= 0.5, without replacement); mean paired difference with 2,000 bootstrap resamples (seed 20260928).",
  "Negative cases: all 44 entities named by all four assistants without an article found automatically were reviewed by Wikipedia search and title lookup (model review)."
 ],
 "limitations": [
  "Observational. Entity signals, prominence and independent coverage are strongly correlated (English Wikipedia vs Tranco score r = 0.654), so their separate contributions are estimated with wide intervals; matched pairs found comparable partners for only 42 of 340 entities with an article.",
  "The frame contains only options at least one assistant named, so the study compares named brands with each other; brands no assistant named are not observed.",
  "Prominence proxies are partial: Tranco measures website traffic, Wikipedia mentions measure encyclopedic reference; news coverage, search demand and revenue are not measured.",
  "Independent coverage is counted inside the pages the assistants cited, fetched two days after the answers, so it is partly the same process as the outcome.",
  "One run per assistant (stability only for 20 questions and three assistants), one date, US English, and only recommendation questions: informational and transactional intents are not covered.",
  "Automated entity matching still misses articles under ambiguous names (8 of the 44 reviewed consensus cases); model review, not human review."
 ],
 "update": "Quarterly, with the next four-assistant run. Planned: repeated runs over several weeks for all four assistants, news and search-demand controls, and controlled tests that change a brand's entity record.",
 "@context": "https://schema.org",
 "@type": "Dataset",
 "name": "Do Wikipedia and schema make AI assistants recommend a brand?",
 "version": "2.0",
 "dateCreated": "2026-09-26",
 "creator": {
  "@type": "Organization",
  "name": "Underneath",
  "url": "https://underneath.agency"
 },
 "url": "https://underneath.agency/research/brand-entity-ai-recommendations-study",
 "temporalCoverage": "2026-09-26/2026-09-28",
 "isAccessibleForFree": true,
 "license": "https://creativecommons.org/licenses/by/4.0/",
 "code": "cite/pipeline/ (collection, validation and analysis scripts)",
 "variableMeasured": [
  "version",
  "controls",
  "reanalysis_of",
  "panel",
  "coverage",
  "wording_types",
  "hierarchy",
  "article_by_assistants_naming",
  "article_measures",
  "models",
  "glmm",
  "interactions",
  "by_wording_national",
  "by_wording_national_n_questions",
  "engines_national",
  "engine_recommend_rates_national",
  "rank_national",
  "own_site_cited_national",
  "pipeline_national",
  "pipeline_local",
  "tiers_national",
  "matched_pairs_national",
  "stability",
  "negative_cases",
  "identity_errors"
 ],
 "dateModified": "2026-09-28",
 "research_questions": {
  "RQ1": "How consistently do ChatGPT, Gemini, Perplexity and Claude recommend the same brands for the same buyer question?",
  "RQ2": "Are entity-representation signals associated with being recommended, at the level of brand x question x assistant?",
  "RQ3": "Does the association remain after accounting for brand prominence (Tranco rank, Wikipedia mentions) and independent coverage in the pages the assistants cited?",
  "RQ4": "Does it differ by assistant, by question wording, and between national and local questions?",
  "RQ5": "Are brands with an article named more consistently across repeated runs? (study 8: 20 questions, 3 assistants, 5 runs)",
  "RQ6": "What do the exceptions look like: brands every assistant named without an article, and brands with an article only one assistant named?"
 },
 "framework": "Proposed, not validated: public web -> entity existence and coherence -> crawl, index, retrieval -> source selection -> generation (recommendation, rank) -> cross-assistant agreement -> stability. This study observes entity signals, prominence proxies, the pages each assistant cited, mention, recommendation, rank, agreement and (on a subset) stability. Crawling, indexing, retrieval beyond citations and user decisions are not observed."
}