{
 "question": "How often do the pages ChatGPT, Gemini, Perplexity and Claude cite rank on Google for the same question, and how much of the gap appears when Google is queried with the searches the assistants ran instead?",
 "design": "Observational, cross-system comparison at one point in time. Not a causal test: Google rankings were not manipulated.",
 "population": "Sources cited by the four assistants for the 80 US buyer questions of study 7 on 26 September 2026 (one run per assistant per question).",
 "inclusion": "Every cited address after normalization; answers without citations count in answer-level denominators of all 80 questions.",
 "exclusion": "Bing results were collected and discarded (the provider returned unrelated pages). This says nothing about whether any assistant uses Bing.",
 "calculations": [
  "URL match after normalization; site match on registrable domain.",
  "Pathways, first match wins: A top 10 for the question as typed (2026-09-26); B top 10 for one of the assistant\u2019s searches; C Google top 100 for the question or a search (2026-09-28); D none of these.",
  "Claude: the searches recorded in the same answers. ChatGPT and Gemini: searches from a separate API run on the same questions and day (proxy). Perplexity: none available.",
  "Rank-weighted score: mean of 1/log2(rank+1) over an answer\u2019s citations, rank in Google\u2019s top 100 for the question (2026-09-28), 0 if absent; times 100.",
  "95% intervals: cluster bootstrap resampling questions (10,000 resamples, seed 23). Engine differences paired by question.",
  "Candidate model: logistic GEE on Google\u2019s top-10 pages x engine, robust errors clustered by question."
 ],
 "limitations": [
  "One run per assistant per question, on one day; run-to-run variation is not measured here.",
  "Search queries for ChatGPT and Gemini come from API models, not the consumer apps whose citations are analyzed.",
  "Google top 100 was collected two days after the answers.",
  "Locality subgroups are small (24 local questions)."
 ],
 "protocol": "cite/experiments/s23-reformulation/config.json (frozen before the version-2 collection)",
 "update": "quarterly",
 "deviations": [
  "1. **Partial SERP results kept (collection).** DataForSEO returned status 40106 (\"Task completed with partial results. Some pages could not be retrieved... You have not been charged for the pages that were not returned\") for many depth-100 requests. The first collection run treated 40106 as a failure and discarded those responses; it was stopped, fixed to keep them (status recorded in each raw file's `meta.task_status`), and restarted. Coverage (share partial, deepest rank reached) is reported in `stats.json` \u2192 `top100_coverage`. Consequence: pathway C (\"top 100\") is a lower bound.",
  "2. **Rank-weighted score date (clarification).** The protocol says \"best Google rank of the cited page in the top 100 for the question as typed\" without a date. The only top-100 collection is 2026-09-28, so the score uses it.",
  "3. **Pathway C definition (clarification).** C counts a citation found anywhere in the 2026-09-28 top 100 for the question or one of the assistant's searches and not already in A or B. This includes pages in the 2026-09-28 top 10 for the question that were not in the 2026-09-26 top 10 (Google drift), reported separately as `citations_page_in_top10_google_2026_09_28`.",
  "4. **Answer-level denominators.** Version 1 computed \"answers citing a top-10 page\" over answers that had at least one citation (61 for Claude, 73 for Gemini, 80 for the others). Version 2 reports both that figure and the share of all 80 questions; the latter is primary.",
  "5. **Bootstrap implementation (analysis, no change to the estimand).** The first implementation rebuilt a resampled data frame for every resample and would have taken hours at 10,000 resamples. It was replaced by frequency weights: a question drawn k times in a resample weighs k, which is the same estimator. A 30-resample dry run of the old implementation is not used anywhere.",
  "6. **One search without results (collection).** 576 of 577 distinct keywords were collected. The remaining one, a ChatGPT proxy search restricted with \"site:\" (`site:insurance.bcbsma.com physical therapy Boston preferred providers`), returned \"No Search Results\" (40102) from the standard queue; it is listed in `stats.json` \u2192 `keywords_missing` and treated as ranking nothing.",
  "7. **Queue bookkeeping file excluded (analysis).** `data/raw/wave2/s23b_queue.json` (posted task ids) matched the raw-file pattern and is skipped when loading results."
 ],
 "stage_model": {
  "search_activation": "whether an answer cites anything",
  "query_rewriting": "Claude's own searches; proxies for ChatGPT and Gemini",
  "retrieval": "Google ranking as a stand-in for an unobserved index",
  "selection_and_citation": "which pages are cited and in what order",
  "use_in_answer": "not measured"
 },
 "@context": "https://schema.org",
 "@type": "Dataset",
 "name": "Do ChatGPT, Gemini, Perplexity and Claude cite pages that rank?",
 "version": "2.0",
 "dateCreated": "2026-09-26",
 "dateModified": "2026-09-28",
 "creator": {
  "@type": "Organization",
  "name": "Underneath",
  "url": "https://underneath.agency"
 },
 "url": "https://underneath.agency/research/ai-citations-google-rankings-study",
 "temporalCoverage": "2026-09-26/2026-09-28",
 "isAccessibleForFree": true,
 "license": "https://creativecommons.org/licenses/by/4.0/",
 "code": "cite/pipeline/s23_google.py (version 1), s23b_collect.py and s23b_analyze.py (version 2)",
 "variableMeasured": [
  "version",
  "answers_collected",
  "google_question_top10_collected",
  "google_top100_collected",
  "questions",
  "citations",
  "keywords_collected_top100",
  "keywords_missing",
  "search_query_counts",
  "bootstrap",
  "top100_coverage",
  "google_drift_top10_2026_09_26_vs_2026_09_28",
  "by_engine",
  "engine_differences_points",
  "candidate_model",
  "locality_counts",
  "reference_ai_overviews_url_in_top10_pct",
  "bing_discarded"
 ]
}