The short version
- Of 269 numbered “best X” lists with an identifiable publisher cited by the six AI surfaces, 65 (24.2%, 95% interval 18.1% to 30.5%) ranked their own publisher first. The denominator is cited numbered lists, not all citations. Our detector misses about one real case in three, so the true share is likely higher.
- Publishers that included themselves almost always went first: 92.9% of self-including lists put the publisher at number one.
- Measured against everything the answers cited, self-ranking lists are a small part: 92 of 8,565 citations (1.1%), and 66 of 1,206 answers (5.5%) cited at least one.
- We found no statistically detectable difference across the six surfaces in this sample (p = 0.41). Shares ranged from 5.9% (ChatGPT, 17 lists) to 34.2% (AI Overviews, 38 lists) with wide intervals, so this does not show that the surfaces behave the same.
- Self-ranking lists are also present in conventional search results: 29.0% of numbered lists in Google’s top 10 for the AI Overview and AI Mode keywords, and 35.2% for the assistants’ questions. These describe different pools of pages from the 24.2% above, not one common share, and the study cannot see which pages the AI engines retrieved.
- Within the 55 answers that cited both a self-ranking and an independent list, there was no statistically detectable difference in this sample in how often each list’s #1 entry was named (odds ratio 1.01, interval 0.52 to 2.75). Across all 590 answer–list pairs, which are not independent, the rates were 47.1% (of 136) and 43.4% (of 454). This is a proxy for answer-level representation, not a measure of influence.
- Within one session, exposure was stable: over five runs minutes apart on ChatGPT, Gemini and Perplexity, in 54 of 60 question-and-surface combinations no self-ranking list was cited in any run, and in 5 one was cited every time. Stability across days was not tested.
What we measured
We took every page cited by the six AI surfaces in our studies of AI Overviews, AI Mode and the four assistants, all US and collected on 26 September 2026, and kept those whose title or address marked them as a ranked list (“best …”, “top 10 …”): 1,574 pages. We fetched each page and read its numbered entries.
A list ranks its publisher first when the publisher’s name begins its first numbered entry. The publisher’s name is the name part of its web address (“rippling” for rippling.com) or the site name the page states about itself.
The six surfaces are not the same kind of system. AI Overviews and AI Mode were captured from Google’s results pages, ChatGPT and Gemini from their consumer apps, and Perplexity and Claude through their APIs (Perplexity’s sonar model, and Anthropic’s claude-haiku-4-5 model with web search). We report them side by side but label the difference.
What this study observes, and what it does not
Visibility in generative search runs through several stages that the research literature treats separately: whether a page is retrieved, what the model is shown, whether the page is cited, how prominently, how much of the answer it shapes, whether the answer represents it accurately, and what the reader does next. This study observes some of these directly, measures others through a proxy, and cannot see the rest.
| Stage | In this study | How |
|---|---|---|
| Retrieval by the AI engine | Not observed | The engines do not expose which pages they retrieved |
| Presence in conventional search results | Observed (Google top 10; Bing not used) | Shows what conventional search offered for the same query; not a measure of AI retrieval |
| What the model was shown, and in what order | Not observed | – |
| Citation | Observed | The list’s address appears among the answer’s sources |
| Prominence | Partial proxy | Whether the list’s #1 was the answer’s first recommendation (assistant answers only); citation order used only as a model covariate |
| Answer-level representation | Proxy | The answer names the list’s #1 entry |
| Absorption (how much of the answer’s facts, wording or structure comes from the source) | Not measured | – |
| Accuracy of what the answer attributes to the source | Not measured | – |
| Exposure across repeated runs | Observed, within one session | Five runs minutes apart on ChatGPT, Gemini and Perplexity |
| Clicks, decisions, conversions | Outside the scope | – |
The study therefore measures one source class at three observable points (citation, conventional-search availability and answer-level representation), with a partial proxy for prominence, plus repeated-run exposure. It is not a model of the whole pipeline, and none of its comparisons identify causes.
How reliable the measurement is
We drew 299 pages across nine groups (self-ranking lists, other numbered lists, pages where parsing rules disagree, unnumbered pages, and cited pages our filter did not flag as lists). Two AI models, Claude Opus and Claude Sonnet, coded each page independently from a saved copy of the page as fetched, without seeing the detector’s result. An AI analyst (Claude) settled their disagreements by reading the page, and each decision is logged. This is model coding, not human coding.
| Check | Result |
|---|---|
| Coders agree the publisher is ranked first | Cohen’s kappa 0.99 |
| Coders agree the page is a ranked list | Cohen’s kappa 0.91 |
| Detector finds real self-first lists (test half) | 64.9% (interval 45.4% to 80.4%) |
| Detector avoids false self-first calls (test half) | 98.0% |
| Self-first calls that are correct (test half) | 93.3% |
| A simpler rule (web address only, every numbered heading) on the same test pages | finds 59.0%, avoids 100.0% |
The detector rarely calls a list self-ranking when it is not, but misses about a third of real cases. In the test half, most misses were pages with more than one numbered section where the detector read the wrong one; the others were a very short brand name and a product name that does not start with the brand (“Ascend by Kindsight”). Of the 65 lists behind the headline, 47 were also in the coded sample, and 4 of those were detector errors, for example two VentureSmarter pages that put ZenBusiness first. We did not adjust the headline for this, because the test sample was too small to estimate the miss rate precisely; read the figures as lower bounds.
Two further checks limit what the data can say. Extending the detector to unnumbered pages (tables, plain lists, structured data) did not work well: it picked the right first entry only 38.7% of the time, so we left those pages out. And of 50 cited pages our title filter did not flag as lists, 5 were ranked lists, 2 of them self-ranking, so some lists are outside the study altogether.
Findings
How often the publisher ranks itself first
| Measure | Unit and denominator | Count | Share (95% interval) |
|---|---|---|---|
| Ranks its publisher first | Cited numbered lists with an identifiable publisher | 65 of 269 | 24.2% (18.1% to 30.5%) |
| Includes its publisher | Cited numbered lists with an identifiable publisher | 70 of 269 | 26.0% |
| Puts the publisher first | Lists that include their publisher | 65 of 70 | 92.9% |
| Points to a self-ranking list | All citations in the answers | 92 of 8,565 | 1.1% (0.7% to 1.5%) |
| Cites at least one self-ranking list | All answers, including those with no citations | 66 of 1,206 | 5.5% (3.9% to 7.1%) |
The 269 lists are those with a usable publisher name; 12 more numbered lists had none. The 65 lists came from 55 different publishers (a few of them detector errors, see above); the most frequent were Zendesk and Rippling (3 lists each) and HubSpot, Klaviyo, Huntress, TMetric and CoCountant (2 each); VentureSmarter’s two lists are left out because both were detector errors. Most are software companies writing about their own category; the rest include a roofing company, a law firm, the car rental company Sixt, a heating and cooling contractor and accounting firms.
368 of the 1,574 pages could not be fetched. If the missing pages contain numbered lists at the same rate as the fetched ones, the true share lies between 18.5% and 41.9% in the worst and best cases.
Adding the answers we had already collected for repeat runs, reworded questions and the US arm of our country study gives 358 lists, of which 21.8% ranked their publisher first (16.8% to 27.3%). In the UK, Canada and Australia arms the share was 18.1% of 72 lists.
By the AI surface that cited the list
| AI surface | How it was collected | Numbered lists cited | Publisher ranks itself first (95% interval) | Answers citing one |
|---|---|---|---|---|
| AI Overviews | Google results page | 38 | 34.2% (21.2% to 50.1%) | 4.7% of 486 |
| Claude | API | 35 | 28.6% (16.3% to 45.1%) | 11.3% of 80 |
| AI Mode | Google results page | 26 | 26.9% (13.7% to 46.1%) | 3.0% of 400 |
| Perplexity | API | 161 | 23.6% (17.7% to 30.7%) | 18.8% of 80 |
| Gemini | Consumer app | 31 | 19.4% (9.2% to 36.3%) | 7.5% of 80 |
| ChatGPT | Consumer app | 17 | 5.9% (1.0% to 27.0%) | 1.3% of 80 |
A list cited by more than one surface counts once for each. A test across all six surfaces (logistic model, errors clustered by publisher) found no statistically detectable difference in this sample (p = 0.41). With 17 to 161 lists per surface, moderate differences could go undetected, so this is not evidence that the surfaces behave the same. Comparisons within the same kind of collection also include zero: AI Overviews minus AI Mode 7.3 points (−10.5 to 25.6), ChatGPT minus Gemini −13.5 points (−31.0 to 5.5), Perplexity minus Claude −5.0 points (−21.9 to 12.0). We therefore do not claim that any surface favors these lists.
Compared with what search shows for the same questions
We compared the cited lists with the numbered lists that Google showed in its top 10 for the same keywords and questions. The top 10 shows what conventional search offered for the same query. It is not a measure of which pages an AI engine retrieved or could have used, which the study cannot see. We used Google only: the Bing results collected for the assistants’ questions were unrelated to the questions and had already been discarded by the study that collected them.
- Among lists in the search top 10, 29.0% (93 lists, AI Overview and AI Mode keywords; interval 16.4% to 43.5%) and 35.2% (54 lists, the assistants’ questions; 22.2% to 48.1%) ranked their publisher first. These are descriptive figures for different source pools from the 24.2% among cited lists, not estimates of a common population share; the formal comparisons are the two below.
- Overlap was low for both kinds of list. 18.6% of 97 citations of self-ranking lists and 11.3% of 300 citations of independent lists (counted per question and surface, including repeat runs and the US arm of the country study) pointed to a page in the same question’s search top 10, a difference of 7.2 points computed before rounding (−3.3 to 19.7), not statistically detectable in this sample.
- Among numbered lists in the same question’s top 10, there was no statistically detectable association between self-ranking and being cited in this sample (odds ratio 0.37, interval 0.04 to 1.54, from only 18 question-and-surface combinations with both kinds). This describes citation among conventional-search results, not how engines select pages.
Self-ranking lists were present both among conventional-search results and among AI citations, but the study cannot observe the AI engines’ internal retrieval process. These are observational associations with conventional-search availability, not evidence about how the engines retrieve or select pages.
Answer-level representation: do answers name the list’s top pick?
For every answer that cited a numbered list (590 answer-and-list pairs in 319 answers, from the main answers plus the repeat runs and the US arm of the country study), we checked whether the answer text names the list’s #1 entry. This measure operationalizes an observable proxy for answer-level representation; it does not directly measure absorption in the broader sense used in the research literature (how much of an answer’s facts, wording or structure comes from a source), and it does not show that the list caused the answer to name the company. Names that appear only as a citation label were not counted. On 120 answer-name pairs coded by two AI models, the rule never counted a name the coders said was absent (100.0%) and found 87.1% of real mentions.
| Kind of list | Answer names the list’s #1 entry |
|---|---|
| Publisher ranked itself first | 47.1% of 136 |
| Independent list | 43.4% of 454 |
The pairs are not independent: one answer can cite several lists, and one question has answers from several surfaces and runs. The pooled difference, 3.7 points (−10.9 to 21.5, resampling questions), is not statistically detectable in this sample. The primary comparison uses only the 55 answers that cited both a self-ranking and an independent list and compares the lists within each answer, which holds the question, surface and answer fixed: odds ratio 1.01, with an interval of 0.52 to 2.75 when the 28 questions behind these answers are resampled (0.59 to 1.75 if the answers were treated as independent). In 16 of the 55 answers only the self-ranking list’s #1 was named, in 19 only the independent list’s #1, in 17 both and in 3 neither (an answer counts as naming a kind if it named the #1 of any cited list of that kind; the odds ratio uses all 211 answer–list pairs in these answers). Among other checks, a model adjusting for surface, industry, citation order, whether the entity has a Wikipedia article, how many cited lists put it first and whether another page from the same publisher was cited, with errors clustered by question, gives 1.36 (0.79 to 2.36). Restricting to #1 entries that clearly read as names reverses the direction (47.4% of 135 against 53.1% of 369 pairs; difference −20.4 to 13.3 points). In the ChatGPT, Gemini, Perplexity and Claude answers, the list’s #1 was the answer’s first recommendation 9.8% of the time for self-ranking lists and 18.6% for independent lists (difference −20.4 to 6.6 points).
The pooled figure hides differences by surface in small samples: in AI Overviews and AI Mode answers the self-ranking publisher was named more often (62.5% of 24 and 61.5% of 13, against 30.4% and 25.8% for independent lists), and in Claude answers less often (50.0% of 10 against 72.0% of 25). These splits are small and were not tested separately.
Other answers to the same question that did not cite the list named the same #1 entries in 26.2% (self-ranking) and 33.4% (independent) of cases.
How stable it is
We had asked 20 questions five times each on ChatGPT, Gemini and Perplexity, minutes apart, so this is stability within one session. In 54 of the 60 question-and-surface combinations no self-ranking list was cited in any run, in 5 one was cited every time, and only 1 changed between runs. Which question it was explained most of the variation (intraclass correlation 0.77). Rewording the question changed whether a self-ranking list was cited in 2 to 5 of 60 pairs, depending on the wording. We have not yet repeated the collection on a different date.
Robustness
Across the choices below, the share stays between 18.6% and 26.7%.
| Variant | Lists | Share ranking publisher first |
|---|---|---|
| Detector used for the headline | 269 | 24.2% |
| Simpler detector (web address only, every numbered heading) | 285 | 18.6% |
| Without pages fetched on a retry | 265 | 24.2% |
| Including Google question keywords | 272 | 23.9% |
| Leaving out the five most frequent self-ranking publishers (ties broken alphabetically) | 253 | 20.9% |
| One list per publisher | 195 | 26.7% |
Reading unnumbered pages as well gives 12.1% of 810 lists, but that reading failed validation (see above), so it is not used.
How this compares with other studies
| Source | Sample and date | Figure |
|---|---|---|
| This study | 269 numbered lists with an identifiable publisher cited by 6 AI surfaces, 8 industries, September 2026 | 24.2% ranked their publisher first; 1.1% of all citations |
| Peec AI | 232,000 citations for software-review prompts, December 2025 to February 2026 | “Approximately 1 in 10 citations” came from self-promotional listicles; ChatGPT 3.6%, AI Mode 10.3%, Perplexity 10.4% |
| Ahrefs | 26,283 ChatGPT source URLs, December 2025 | “Best X” lists were 43.8% of cited page types |
| Lily Ray | 100 B2B queries in AI Overviews, April to June 2026 | When a brand’s own list was cited, the brand “was left out of the actual recommendation 69% of the time” |
| Scrunch | About 10,000 cited URLs, May to June 2026 | Reports that brands cited via their own list had a recommendation rate of “roughly 7%”, against “about 4%” |
Peec AI measures self-promotional lists as a share of all citations in software prompts; our comparable figure across eight industries is 1.1%. Lily Ray found that brands cited through their own lists were usually left out of the recommendation, and Scrunch reported a small difference in recommendation rate; both measure recommendation. We measure whether answers named the list’s top pick, and found no statistically detectable difference in this sample. The studies use different samples and measures.
Sources: Peec AI (opens in a new tab); Ahrefs (opens in a new tab); Lily Ray (opens in a new tab); Scrunch (opens in a new tab) (figures checked via PPC Land (opens in a new tab)).
How this relates to research on generative engine optimization
- Aggarwal et al. (2024), “GEO: Generative Engine Optimization” (arXiv (opens in a new tab)), showed that rewriting a page’s content changed how much of a generated answer was attributed to it, in a setting where the page was already among the few sources given to the model. It studied content changes mainly with retrieval held fixed (the top five Google results placed in the model’s context), plus a smaller test on Perplexity.
- Chen et al. (2025), “Generative Engine Optimization: How to Dominate AI Search” (arXiv (opens in a new tab)), found that AI search engines drew on different source ecosystems from Google and from one another, citing earned media more heavily than brand-owned and social sources, with differences across engines, languages and phrasings. The authors note that, in their classification, the line between an earned outlet and a brand’s own blog can sometimes be blurry.
- Martinez (2026), a critical survey of the field (arXiv (opens in a new tab)), argues that visibility in generative search is a multistage, partly unobservable and variable process, and that retrieval, citation, prominence, absorption, accuracy and downstream behavior should be measured separately.
- This study applies that stage-by-stage view to one source class that sits on the blurry line Chen et al. describe: a brand-owned page written as a comparative review. It separates citation, conventional-search availability, answer-level representation and repeated-run exposure, without observing the engines’ retrieval or identifying causes.
What this means
These points separate what was measured from interpretation, what is not established, practical implications and open questions.
- Measured: self-ranking commercial lists are a measurable subset of the sources cited by these six surfaces on 26 September 2026 (US): about a quarter of cited numbered lists with an identifiable publisher, and 1.1% of all citations.
- Measured: in this sample, answers named a self-ranking list’s own publisher at a rate not statistically distinguishable from the top pick of an independent list. Ranking oneself first was not associated with a statistically detectable difference in being named in this sample (within-answer odds ratio 1.01, 0.52 to 2.75); the design cannot say whether self-ranking has any effect.
- Interpretation: readers of AI answers that cite such lists are sometimes reading a vendor’s ranking of itself.
- Not established: whether self-ranking lists improve or harm readers’ decisions.
- Practical implication (not a measured finding): a list that states its publisher’s interest lets readers weigh it accordingly.
- Open question: in small samples, answers in Google’s AI surfaces named a self-ranking publisher more often than an independent list’s top pick; a larger study should test this directly.
Methodology
- Pages: every page cited by AI Overviews and AI Mode (800 US head and tail keywords) and by ChatGPT, Gemini, Perplexity and Claude (80 US buyer questions), 26 September 2026, whose title or address contains “best”, “top N”, “top rated” or “N best”; platform pages (YouTube, Reddit, social networks, Google, Amazon, Yelp, Wikipedia) excluded; 1,574 pages. Secondary sets: 20 questions repeated five times, three rewordings of 20 questions, and 40 questions in the US, UK, Canada and Australia.
- Fetching: plain HTTP, then a rendering service (Firecrawl) for every page that failed; 1,206 of the 1,574 pages fetched.
- List entries: the longest run of consecutive numbered H2 or H3 headings (“1.”, “2.”, “3.”), at least three entries.
- Publisher: the name part of the registrable domain, plus the site name stated on the page; names under four letters or generic words excluded.
- Validation: 299 pages coded by two AI models, split in half to choose and then test the rule; agreement, adjudication log and results in the data files.
- Search comparison: Google’s organic top 10 for the same keyword (AI Overviews, AI Mode) and for the same question (assistants); Bing results were not used.
- Answer mentions: whole-name match of the list’s #1 entry in the answer text, after removing citation labels.
- Units of analysis: the data nest as question → surface and run → answer → cited address → page and list → publisher → list entry. List shares use one row per cited list (clustered by publisher); per-surface shares use one row per list and surface; citation shares use citations and answer shares use answers (both clustered by question); the search comparison uses citations per question and surface, and lists within each question-and-surface pair; answer-level representation uses answer–list pairs, with the primary comparison within answers; stability uses question-and-surface combinations across runs.
- Statistics: 95% intervals by bootstrap with 2,000 resamples of publishers (headline shares) or questions (citations, answers), and by Wilson’s method (exact at zero) for per-surface shares; surface differences tested with a logistic model. The analysis was not pre-registered; decision rules were written down before the coding, and every deviation was logged.
- Update schedule: a second collection date, with repeated runs and reworded questions, is planned for October 2026.
Limitations
- The detector misses about a third of real self-ranking lists, and 368 of the 1,574 pages could not be fetched, so the shares are likely undercounts.
- Only numbered lists are measured; unnumbered round-ups could not be read reliably, and the title filter misses some lists.
- The validation sample was coded by two AI models, not by people.
- Google’s top 10 shows what conventional search offered for the same query; it is not a measure of which pages the AI engines retrieved or could have used, which the study cannot see.
- Answer mentions show what answers say, not why they say it, and are only a proxy for representation. The #1 entry of a self-ranking list is a brand name the detector matched, while an independent list’s #1 is read less reliably, so measurement error may differ between the two groups; restricting to #1 entries that clearly read as names reverses the small difference.
- All data comes from one collection date; the repeated runs were minutes apart, and per-surface samples are small.
Data and downloads
- Every cited list page with its fetch status, class, first entry and self-rank: s13_list_pages.csv and JSON
- The 299-page gold standard with both coders’ labels, the codebook and the adjudication log: gold_standard.csv, codebook, adjudication.csv
- Every statistic on this page: stats.json
- Machine-readable methodology: methodology.json
- The decision rules written before the coding, with every deviation: analysis_plan.txt
The data is free to reuse with attribution (CC BY 4.0). To cite: Underneath (2026), How many “best of” lists cited by AI rank their own brand first?, Underneath Research, https://underneath.agency/research/self-promoting-best-lists-study
Frequently asked questions
Do AI engines cite self-promotional “best of” lists?
Yes. Of 269 numbered “best X” lists with an identifiable publisher cited by AI Overviews, AI Mode, ChatGPT, Gemini, Perplexity and Claude, 24.2% were published by a company that ranked itself first. They made up 1.1% of all citations.
Which AI engine cites self-ranking lists most?
We found no statistically detectable difference in this sample. Shares ranged from 5.9% for ChatGPT to 34.2% for AI Overviews, but the samples are small, so this does not mean the surfaces behave the same.
Are self-ranking publishers named more often in answers that cite their list?
Not detectably in this sample. Within the 55 answers that cited both kinds of list, there was no statistically detectable difference in how often each list’s #1 entry was named (odds ratio 1.01, interval 0.52 to 2.75). Across all answer–list pairs, which are not independent, the rates were 47.1% (self-ranking) and 43.4% (independent). This is an observational comparison of answer-level representation, not a test of influence.
How common is it for a list publisher to include itself?
26.0% of cited numbered lists with an identifiable publisher included the publisher somewhere, and when they did, 92.9% ranked it first.