They are. The assistants write their searches in different ways. When a search names a particular source, that source is cited far more often than when it is not named. But named sources explain only a small part of what gets cited. Most cited pages come from ordinary searches that name no source.
The short version
- ChatGPT ran a mean of 3.7 searches per answer (95% interval 3.1 to 4.2), Gemini 1.9 and Claude 0.76. None of the 509 searches repeated the user’s question word for word. Every answer that ran a search cited at least one page, and none of the 42 answers without a search cited anything.
- The assistants search for different things. In 43.8% of its answers ChatGPT ran a search aimed at a named publication, ranking or award, against 26.2% for Gemini and 1.2% for Claude. ChatGPT also looked for reviews in 46.2% of answers and prices in 23.8%. Claude’s searches were almost all plain searches for options.
- When a search named a source, such as NerdWallet, Avvo or Super Lawyers, the answer cited that source 44.0% of the time (40 of 91). In answers from the same assistant and industry that did not name it, the figure was 8.1%. That is an association, not proof that naming causes citing.
- Only 4.2% of cited domains (40 of 943) had been named in a search of the same answer. Gemini cited Reddit in 31 answers but mentioned Reddit in a search in only 5.
- More searches go with more cited domains: with twice as many searches, an answer cited 1.45 times as many domains, after allowing for assistant and industry. At the same number of searches, Gemini cited 3.29 times as many domains as ChatGPT.
- Version 1.0 reported that 10.5% of ChatGPT’s searches named a past year. Most of those looked for the latest edition of an annual ranking, such as J.D. Power 2025. Only 3 of ChatGPT’s 296 searches used an old year for no apparent reason, and Gemini’s past years always came with 2026.
Research questions
- RQ1. How many searches does each assistant run, and how far do they move from the question?
- RQ2. What do the searches try to find: options, current information, named authorities, reviews, prices, comparisons, particular platforms?
- RQ3. Do answers that run more searches cite a wider range of websites, for the same assistant and industry?
- RQ4. When a search names a source, is that source cited in the same answer more often than when it is not named?
- RQ5. Are the searches, and their link to citations, stable across repeated runs, dates and wordings? This is not answered here: we have one run per question and assistant.
What we analyzed
For each of the 80 buyer questions (10 in each of eight industries, such as “What is the best CRM for a small business?” and “Who is the best divorce lawyer in Houston?”), we asked ChatGPT (GPT-5.4 nano) and Gemini (Gemini 3.5 Flash-Lite) through their APIs with web search on. We recorded the searches each one reported running and the pages it cited. Claude’s searches and citations (Claude Haiku 4.5, also with web search) come from the answers collected for the four-assistant study the same day. That gives 240 answers and 509 searches. Version 1.1 re-analyzes the same answers; no assistant was asked anything new.
Each search was analyzed in three ways. Fixed phrase rules detect a year, review words, a site: operator and similar patterns, as in version 1.0. A model coder labeled what each search is for and the role of any year in it, without being told which assistant ran it. And a sentence-embedding model measured how close each search is in meaning to the original question.
Finding 1: how many searches, and how far from the question
| Measure | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Answers with a search | 73.8% | 97.5% | 76.2% |
| Mean searches per answer | 3.7 | 1.9 | 0.76 |
| 95% interval | 3.1 to 4.2 | 1.8 to 2.0 | 0.7 to 0.8 |
| Total searches | 296 | 152 | 61 |
| Median words per search | 8 | 6 | 6 |
| Median similarity to question | 0.73 | 0.85 | 0.9 |
| Search functions per answer | 5.76 | 3.35 | 2.36 |
Similarity runs from 0 to 1; 1 means the search says the same thing as the question. ChatGPT moved furthest from the question: it added a median of 5 words that were not in it, Gemini 2 and Claude 1. Claude never ran more than one search per answer. The last row counts the distinct functions (see the next section) covered by an answer’s searches.
Intervals come from resampling the 80 questions, so answers to the same question are not treated as independent.
Finding 2: what the searches are for
A model coder read every search next to its question and recorded each function the search performs. The coder was not told which assistant ran it.
| Share of answers with a search that… | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Looks for options | 72.5% | 97.5% | 76.2% |
| Asks for current information | 52.5% | 76.2% | 43.8% |
| Looks for a specific feature | 47.5% | 30.0% | 20.0% |
| Looks for reviews or ratings | 46.2% | 30.0% | 5.0% |
| Targets a named authority | 43.8% | 26.2% | 1.2% |
| Restricts to a place | 40.0% | 50.0% | 30.0% |
| Compares options | 33.8% | 0.0% | 0.0% |
| Targets a platform or directory | 31.2% | 10.0% | 0.0% |
| Looks for prices | 23.8% | 6.2% | 3.8% |
| Targets an option’s own site | 23.8% | 0.0% | 0.0% |
| Looks for complaints | 6.2% | 0.0% | 0.0% |
“Named authority” means a publication, ranking or award, such as U.S. News, Wirecutter, J.D. Power or Super Lawyers. “Platform or directory” means a site such as Avvo, Yelp, G2 or Reddit, often searched with a site: operator. A search can do several things at once, so the rows do not add up to 100%.
The phrase rules of version 1.0 give the same picture. With 95% intervals: 37.5% of ChatGPT’s answers had a search naming a publication or directory (27.5% to 47.5%), 41.2% added review words (30.0% to 52.5%) and 10.0% used a site: operator (3.8% to 17.5%).
ChatGPT’s searches split one question into several jobs. Its most common primary purpose was finding options (32.4% of its searches), followed by reaching a named authority (23.0%). For Gemini, 86.8% of searches were primarily about finding options, and for Claude 98.4%.
What a real set of searches looks like
For “Who is the best divorce lawyer in Houston?”, ChatGPT ran four searches:
- best divorce lawyer in Houston TX AVVO top rated family law attorney
- Houston divorce attorney Martindale-Hubbell top rated family law
- Houston family law divorce attorney U.S. News Best Lawyers Houston divorce
- Super Lawyers Houston family law divorce 2025 2026
Gemini ran two (“best divorce lawyers in Houston Texas” and “top rated family law attorneys Houston TX”). For a question about San Francisco employment law firms, ChatGPT searched inside single directories: “site:lawyers.com employment law San Francisco” and “site:avvo.com employment lawyer San Francisco review”.
Finding 3: years in the searches
Assistants write the year into their searches: 52.5% of ChatGPT’s answers, 76.2% of Gemini’s and 43.8% of Claude’s included a search with a year. Version 1.0 counted any year before 2026 as out of date. The model coder instead recorded why the year was there.
| Year in the search | ChatGPT | Gemini | Claude |
|---|---|---|---|
| No year | 170 | 81 | 26 |
| 2026 only | 95 | 46 | 35 |
| 2026 and an earlier year | 14 | 25 | 0 |
| Earlier edition of a ranking | 14 | 0 | 0 |
| Earlier year, no reason | 3 | 0 | 0 |
Counts are searches. An “earlier edition” search asks for an annual ranking whose latest edition may carry the earlier year, such as “J.D. Power homeowners insurance customer satisfaction Florida 2025” or “Gartner Magic Quadrant 2024 managed detection and response”. Gemini’s past years always came alongside 2026 (“best boutique hotels Nashville 2025 2026”). Only 3 searches, all ChatGPT’s, used an earlier year with no apparent reason.
Finding 4: from searches to citations
A named source is often cited, but most citations are not named
We matched source names in the searches (Avvo, NerdWallet, U.S. News, Reddit and 28 others, listed in the dataset) to their websites. Then we checked whether the answer cited that website.
| Source named in a search | ChatGPT | Gemini | Both |
|---|---|---|---|
| Times named | 68 | 23 | 91 |
| Cited in the same answer | 45.6% | 39.1% | 44.0% |
| Cited when not named | 5.1% | 11.4% | 8.1% |
“Cited when not named” uses answers from the same assistant and the same industry, for the same source, whose searches did not name it. Claude never named a source. In a model that accounts for answers to the same question, the odds of a source being cited were 10.0 times higher when a search named it (95% interval 5.92 to 16.91).
Sources differed. Reddit was cited all 5 times it was named, and NerdWallet 5 times out of 9. U.S. News was named 7 times and never cited, Yelp 6 times and never cited. Of 12 searches inside one site with a site: operator, 5 led to that site being cited.
The reverse view is smaller. Of the 943 cited domains, counted once per answer, 40 (4.2%) had been named in one of that answer’s searches. Gemini cited Reddit in 31 of its 80 answers and mentioned Reddit in a search in 5 of them. Most cited pages arrive through ordinary searches for options, not searches aimed at them.
A broader search purpose did not show the same link. Answers with a search targeting a named authority or platform cited at least one of the listed sources in 55.3% of cases (76 answers), and answers without one in 54.1% (122 answers). The odds ratio was 0.86 (0.39 to 1.9). The link is with the specific source named, not with the general habit of searching for authorities.
More searches, more cited websites
Within each assistant, answers with more searches cited more distinct domains: Spearman correlation 0.35 for ChatGPT and 0.3 for Gemini. Claude ran at most one search, so it cannot show this. In a model that accounts for answers to the same question and for industry, twice as many searches went with 1.45 times as many cited domains. At the same number of searches, Gemini cited 3.29 times as many domains as ChatGPT (2.41 to 4.49) and Claude 2.98 times as many (1.8 to 4.96).
So the number of searches is not the whole story. ChatGPT searched the most and cited the fewest websites: 2.38 domains per answer, against 6.28 for Gemini and 3.14 for Claude. How many results each assistant keeps from a search matters as much as how many searches it runs.
Three ways of searching
| Measure | ChatGPT | Gemini | Claude |
|---|---|---|---|
| Mean searches | 3.7 | 1.9 | 0.76 |
| Similarity to question | 0.73 | 0.85 | 0.9 |
| Targets a named authority | 43.8% | 26.2% | 1.2% |
| Targets a platform | 31.2% | 10.0% | 0.0% |
| Looks for reviews | 46.2% | 30.0% | 5.0% |
| Cited domains per answer | 2.38 | 6.28 | 3.14 |
Shares are of answers; similarity is the median per search.
The three assistants follow three different patterns. ChatGPT runs many narrow searches, aimed at named authorities, directories, reviews and prices, and cites few websites. Gemini runs one or two broad searches, almost always with the year, and cites many websites. Claude runs a single search close to the question and cites a moderate number. The patterns describe these API models on one day, not a fixed trait of each company’s product.
Observed, inferred and unknown
- Observed. The searches each API reported, the pages each answer cited, and how the two line up within an answer.
- Inferred. Naming a source in a search is associated with citing it. The link is strong and consistent across both assistants that name sources, but engines, questions and searches were not randomized. Questions that lead an assistant to name Avvo may also be questions where Avvo would be cited anyway.
- Unknown. Which results each search returned, and therefore whether a page was found and passed over, or never found at all. Whether the answer’s claims match the cited pages. Whether the same question gives the same searches tomorrow. Whether being cited leads to clicks or customers.
A source can be found but unused, used but uncited, cited but unsupported, visible one day and gone the next, or visible without anyone acting on it. These are different outcomes. This study measures the first step and the last, and only for one run.
What this means
This section is our interpretation.
- Being in the sources an assistant names is a direct route. When ChatGPT names Avvo, Super Lawyers or NerdWallet in its own searches, those sites are cited far more often. A business listed there, with a good profile, sits on that route. The route is narrow: most citations come from searches that name no source.
- Different assistants reward different presence. ChatGPT’s searches reach for rankings, directories and reviews. Gemini’s broad searches bring back many pages, often from Reddit and YouTube. Claude’s single search returns what a plain search for options returns. One tactic does not cover all three.
- Freshness is written into the query, often through rankings. Many searches ask for the current year. ChatGPT often asks for the latest edition of an annual list, which favors pages that are clearly dated and kept current.
Where this sits in GEO research
Early work on generative engine optimization showed that changing a page can change how often an AI answer uses it, once the page is in front of the model. Later work compared which sources different AI search systems cite. More recent reviews argue that “visibility” mixes several outcomes: being found, being used, being credited, staying visible, and producing value. They also note that much of the evidence covers only the step after a page is found.
This study looks at the step before: the searches that decide what can be found. It links those searches to the final citations, at the level of the answer. The next questions need other designs. The first is to record which results each search returned, to separate “never found” from “found and passed over”. The second is to repeat the same questions over several days to test stability. The third is to check claims against cited pages. The fourth is to add or remove a source under controlled conditions, to test cause rather than association.
Methodology
- Questions: the 80 US buyer questions of our four-assistant study, 10 per industry.
- Engines: ChatGPT (GPT-5.4 nano) and Gemini (Gemini 3.5 Flash-Lite) through DataForSEO’s LLM Responses API with web search, 26 September 2026. Claude (Haiku 4.5) through the same API, collected for the four-assistant study the same day. The searches are those each API reported (“fan-out queries”); cited pages are the answer’s annotations, reduced to registrable domains and counted once per answer.
- Phrase rules: fixed rules applied to every search (listed in the analysis code), unchanged from version 1.0.
- Function coding: Claude Opus coded every search, with the question shown, the assistant hidden and the order shuffled. It recorded every function from a list of 12, the primary function, and the role of any year. Claude Sonnet coded the same searches independently. They agreed on the primary function for 78.6% of searches (kappa 0.69) and on the year role for 99.4% (kappa 0.99). The median kappa across the 12 functions was 0.9; the weakest were “pins down which company” (0.11) and “looks for a specific feature” (0.66), so read the feature row in Finding 2 as indicative.
- Similarity: cosine similarity between each search and its question, using the all-mpnet-base-v2 sentence-embedding model.
- Named sources: a fixed list of 32 source websites, matched from source names in the searches, set before the citations were examined.
- Models: GEE regression with answers grouped by question: logistic for “named source cited”, with the assistant as a fixed effect; Poisson for distinct cited domains, with assistant, industry and the log of the number of searches.
- Intervals: 95% bootstrap resampling questions (2,000 resamples).
- Update schedule: quarterly.
Limitations
- One answer per question and assistant, on one date. Whether the searches, and their link to citations, hold across repeated runs, days and rewordings is not measured.
- The APIs report the searches but not the results of each search. We cannot tell a page that was never found from one that was found and not used.
- The link between searches and citations is an association. Questions, assistants and searches were not assigned at random.
- These are API models, not the consumer apps; the consumer ChatGPT and Gemini may search differently.
- The function and year labels were coded by two AI models from the same family. Their agreement checks consistency, not accuracy. No person coded the searches.
- The named-source list is fixed, so searches naming other sources are not matched. Whether the answers’ claims are supported by the cited pages was not checked.
What changed in version 1.1
Version 1.0 (26 September 2026) described the searches with phrase rules. Following an external review, version 1.1 re-analyzes the same 240 answers. It places the searches in the steps that lead to a citation and adds the cited pages of each answer. It also adds model-coded search functions and year roles, a similarity measure, intervals and models that group answers by question, and a second coder.
The phrase-rule figures are unchanged. The statement that 10.5% of ChatGPT’s and 16.4% of Gemini’s searches “asked for a past year” is replaced by year roles. Most of those searches asked for the latest edition of an annual ranking or paired the earlier year with 2026. The earlier advice that a business absent from named directories “is less likely to be found” is reworded as an association. The full change log is in the dataset.
Data and downloads
- Every search with its functions, year role, similarity and named sources: s19_searches_v11.csv and JSON
- Every answer with its searches, named sources and cited domains and pages: s19_answers_v11.csv and JSON
- Every statistic on this page, including models and agreement: stats.json
- Version 1.0 searches and statistics: s19_fanout_queries.csv and stats_v1.0.json
- Machine-readable methodology: methodology.json
The data is free to reuse with attribution (CC BY 4.0).
To cite: Underneath. (2026). The hidden searches AI assistants run before they answer (Version 1.1). Underneath Research. https://underneath.agency/research/ai-hidden-searches-study
Frequently asked questions
What is query fan-out?
When an AI assistant uses web search, it writes its own search queries, often several, instead of searching for the user’s question as typed. Those searches decide which pages it can read and cite. In our test ChatGPT ran a mean of 3.7 searches per buyer question, Gemini 1.9 and Claude 0.76.
Does a source named in ChatGPT’s search get cited?
Often. When ChatGPT’s search named a source such as NerdWallet, Avvo or Super Lawyers, the answer cited that source 45.6% of the time, against 5.1% in comparable answers that did not name it. But only 4.2% of all cited domains had been named in a search.
Does ChatGPT search for old years?
Rarely without reason. 52.5% of ChatGPT’s answers included a search with a year, mostly 2026. Searches with an earlier year usually looked for the latest edition of an annual ranking; 3 of its 296 searches used an old year for no clear reason.
Do more searches mean more sources?
Within an assistant, yes: twice as many searches went with 1.45 times as many cited domains. Between assistants, no: ChatGPT searched most but cited the fewest domains per answer.
Do AI assistants search Reddit?
Rarely by name: 6.2% of Gemini’s answers included a search mentioning Reddit, and none of ChatGPT’s or Claude’s did. Gemini still cited Reddit in 31 of its 80 answers, found through ordinary searches.