---
title: "What pages cited by AI Overviews have in common: 3,096 pages"
description: "3,096 top-10 pages on the same searches: rank predicted AI Overview citation. Compared within each search, only a machine-readable date held up."
canonical: "https://underneath.agency/research/ai-overview-cited-pages-study"
published: 2026-09-26
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Research · AI search

# What pages cited by AI Overviews have in common: 3,096 pages

Most advice about getting cited in Google’s AI Overviews compares cited pages with “the web”. We compared them with the pages they beat: for 486 US searches that showed an AI Overview, we fetched every page ranking in the top 10 and compared the ones the AI Overview cited with the ones it passed over on the same results page. Version 1.1 goes further. It compares each page only with its competitors on the same search, enters all 14 page features at once, and follows cited pages past the citation: whether they were the answer’s first source, and how much of the wording that cites them can be found on the page. Ranking position still came first. Of the on-page features, only a machine-readable date held up, and only on informational searches.

## The short version

1. Position came first: 41.7% of pages ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Within the same search, position accounted for 6.1% of the variation in which pages were cited, and the 14 page features added 2.5 points more; 91.4% was left unexplained by anything we measured.
2. Compared only with pages on the same search, with all features entered together, one feature survived a correction for testing 14 at once: a machine-readable date (+7.9 points, 95% interval 2.9 to 12.9).
3. Question-form headings, the largest difference in version 1.0 (+9.6 points), fell to +3.5 (interval −0.3 to 7.1) and are no longer distinguishable from zero. About half of the original gap came from which searches have such pages; entering the other features took off a little more.
4. The date association is confined to informational searches: +17.8 points there, against +1.0 on commercial and transactional searches. Article schema shows the same split (+10.9 against +0.1).
5. Organization schema, BreadcrumbList schema and HTML tables showed no meaningful association with citation; their intervals sit inside ±5 points.
6. Being cited is not the end of the story. Only 21.0% of cited pages were the answer’s first source, and in 62.2% of cited pages no passage citing the page had at least 80% of its content words on the page. Pages with a table were 14.7 points more likely to have their passages absorbed, although tables did nothing for citation itself.

## Why this version separates the stages

An AI Overview does not rank pages; it picks some, then writes with them. A page can rank and not be picked, be picked and appear only as the fifth link, or be linked while the sentence it is attached to says something the page does not. Treating “visibility” as one number hides which of these steps a change to a page affects.

Version 1.1 therefore measures three outcomes separately, each among the pages that reached the previous step.

| Stage | Question | Pages |
|---|---|---|
| 1. Citation | Is a top-10 page cited at all? | 3,096 |
| 2. Prominence | Is a cited page the first source, and what share of passages link it? | 870 |
| 3. Absorption | Is the wording that cites it found on the page? | 842 |

Three steps sit outside this data. Whether Google retrieved a page at all (its candidate set, including pages outside the top 10) is not visible. Whether the answer states faithfully what the page says (fidelity) needs a reading of meaning, not word overlap. What users then do is not observed.

A second separation matters as much. This study measures how often pages with a trait are cited. It does not measure what happens when a trait is added to a page, because no page was changed. A 2023 lab study of generative engine optimization (Aggarwal et al.) reported that rewriting a source could raise its visibility by up to 40%, but in that setup the source was already among the documents given to the engine. That result and ours answer different questions, and neither says whether adding a trait moves a real page into a live AI Overview.

## What we measured

We started from the 486 US searches in our [AI Overview frequency study](https://underneath.agency/research/ai-overviews-frequency-study) that showed an AI Overview, and took every organic result in the top 10, leaving out platform pages (YouTube, Reddit, Facebook, Google’s own pages and other social sites) whose structure the site owner does not control. Each page was labeled cited or not cited according to whether its URL appeared in that search’s AI Overview.

We fetched each page directly, and through Firecrawl when the direct request failed, returned an error status or returned too little HTML; Firecrawl supplied 612 of the usable pages. In all, 3,096 of 3,624 pages (85.4%) returned a usable HTML page: 936 cited and 2,160 not cited. Fetch success was nearly identical for the two groups (86.1% and 85.1%).

For each page we recorded its JSON-LD types, its newest machine-readable date, author signals, word count, headings, question-form headings, FAQ sections, tables and lists.

New in version 1.1, we read the text of each AI Overview, which links its sources passage by passage. Of the 936 cited pages, 870 (92.9%) are linked inside the answer text; the rest appear only in the source list. No uncited page appeared in the text, which confirms the version 1.0 labels.

## Stage 1: which ranking pages get cited

### Position

| Organic position | Cited | Not cited | Share cited |
|---|---|---|---|
| 1 to 3 | 459 | 642 | 41.7% |
| 4 to 6 | 291 | 779 | 27.2% |
| 7 to 10 | 186 | 739 | 20.1% |
| All | 936 | 2,160 | 30.2% |

Within the same search and holding the page features constant, a page at position 2 was 12.9 points less likely to be cited than the page at position 1, and a page at position 9 was 34.1 points less likely.

### On-page features, three ways

The table gives the difference in the chance of being cited, in percentage points, between pages with and without each feature. “1.0” is the version 1.0 estimate (cited against uncited within position bands). “Alone” compares pages on the same search at the same position, one feature at a time. “Together” does the same with all 14 features and word count entered at once.

| Feature | 1.0 | Alone | Together | Interval |
|---|---|---|---|---|
| Machine-readable date | +7.0 | +8.0 | +7.9 | 2.9 to 12.9 |
| VideoObject schema | +3.5 | +18.3 | +12.1 | 2.1 to 22.2 |
| Product schema | +2.9 | +15.3 | +10.1 | 1.9 to 18.2 |
| FAQPage schema | +6.0 | +7.5 | +5.9 | 1.2 to 10.5 |
| Question heading | +9.6 | +4.9 | +3.5 | −0.3 to 7.1 |
| BreadcrumbList schema | +3.6 | +4.5 | +1.3 | −2.3 to 4.8 |
| HTML table | +1.9 | +2.0 | +0.6 | −3.3 to 4.1 |
| Organization schema | +5.9 | +3.9 | +0.5 | −3.5 to 4.7 |
| Review schema | +4.5 | +8.3 | +0.4 | −6.3 to 6.6 |
| Article schema | +5.5 | +2.2 | −1.0 | −6.3 to 4.2 |
| Author signal | +4.8 | +2.1 | −2.1 | −7.4 to 2.9 |
| FAQ heading | +4.6 | +2.4 | −2.5 | −6.9 to 1.8 |
| Dated within 90 days | +1.8 | +1.7 | −4.0 | −8.7 to 0.8 |
| LocalBusiness schema | −2.9 | −2.6 | −6.2 | −11.9 to −0.5 |

The 95% interval is for the “Together” estimate. A question heading is an H2 or H3 phrased as a question; review schema includes ratings; LocalBusiness includes its subtypes. “Dated within 90 days” in the “Together” column is the difference beyond simply having a date.

Five of the 14 “Together” intervals exclude zero. After a Holm correction for testing 14 features, only the machine-readable date remains (adjusted p = 0.028). VideoObject and Product schema are the largest estimates but rest on few pages (6.6% and 9.3% of cited pages) and have wide intervals; neither survives the correction.

### Robustness

The same model was run two other ways. Resampling by website instead of by search, because some domains rank on many searches, widened every interval; the date association was the only one still clear of zero (+7.9, interval 1.3 to 13.4). Dropping the 612 pages fetched through Firecrawl left the date (+6.9) and gave question-form headings a narrow positive interval (+4.7, interval 0.6 to 8.7), which again did not survive the correction. It also made recency beyond having a date slightly negative (−5.6 points, interval −10.9 to −0.7).

### Negative results

Three features showed no meaningful association with citation. Their intervals lie inside ±5 points and they are nowhere near significance.

- Organization schema: +0.5 (interval −3.5 to 4.7)
- BreadcrumbList schema: +1.3 (interval −2.3 to 4.8)
- HTML table: +0.6 (interval −3.3 to 4.1)

Word count did not matter once position and search were held constant: no length quartile differed from the shortest by a clear margin. Recency beyond having a date did not help either; pages dated in the last 90 days were, if anything, 4.0 points less likely to be cited than other dated pages (interval −8.7 to 0.8).

## When a feature matters depends on the search

The keyword database labels each search by intent. Informational searches had cited pages more often (39.3% of their top-10 pages) than commercial and transactional ones (27.1%). The associations also differed.

| Feature | Informational | Commercial | Gap | Interval |
|---|---|---|---|
| Date | +17.8 | +1.0 | 16.8 | 8.1 to 25.8 |
| Article schema | +10.9 | +0.1 | 10.8 | 2.3 to 18.9 |
| Organization | +11.6 | +3.5 | 8.1 | −1.4 to 16.4 |
| Question heading | +9.3 | +10.7 | −1.4 | −10.0 to 7.1 |
| FAQPage | +3.1 | +7.8 | −4.6 | −12.1 to 2.5 |

Commercial includes transactional searches; the interval is for the gap. These use the version 1.0 band-adjusted method within each group (133 informational searches, 321 commercial or transactional). A dated article is a common shape for an informational answer and an unusual one for a product or service search, so the date and Article associations may mark the type of page that fits the question rather than the markup itself.

Position did not change the picture much. Adding a term for “feature on a top-3 page” to the model gave −1.5 points for question-form headings (interval −7.7 to 4.8) and +3.3 for dates (interval −3.3 to 9.3): no evidence that either matters more or less near the top.

## Stage 2: first source or one of many

The median AI Overview in this sample linked 8 distinct sources in its text. Among the 870 cited pages linked in the text, 183 (21.0%) were the first source linked. The median cited page was linked from 25.0% of the answer’s cited passages.

| Organic position | First source | Passage share |
|---|---|---|
| 1 to 3 | 27.5% | 31.3% |
| 4 to 6 | 14.9% | 29.1% |
| 7 to 10 | 15.7% | 32.1% |

Both columns are among cited pages: the share that were the first source, and the mean share of the answer’s passages that link them. A cited top-3 page was 1.8 times as likely as a cited page at 7 to 10 to be the first source, but once cited, lower-ranked pages were linked from as many passages. No page feature predicted being the first source after the correction. The largest estimate, a machine-readable date, pointed the other way (−11.8 points, interval −23.7 to 0.1).

## Stage 3: is the cited wording on the page

For each passage in an AI Overview that links a page, we measured the share of its content words (stop words removed) that appear anywhere in that page’s text. This is a lexical proxy for absorption. It shows whether the answer’s wording can be traced to the page, not whether the answer is faithful to it.

- On 800 cited pages from searches that also had uncited pages to compare, 63.8% of the content words in the passages citing a page were on that page. The same passages scored against the uncited top-10 pages of the same search gave 53.5%. The gap, 10.3 points (interval 8.2 to 12.2), shows the measure tracks the cited source rather than the topic alone.
- Across 842 scored cited pages, 27.0% of citing passages, on average, had at least 80% of their content words on the page. In 62.2% of cited pages, no citing passage reached that level.
- The median share of five-word sequences copied verbatim from the page was 0.0%. AI Overviews paraphrase; they rarely lift sentences.

One feature stood out after correction. Cited pages with an HTML table were 14.7 points more likely to have their citing passages absorbed (interval 6.5 to 22.9, adjusted p = 0.014), and were linked from 5.0 points more of the answer’s passages (interval 1.3 to 9.0, not significant after correction). Tables showed no association with being cited at all (+0.6). A machine-readable date went the other way, with less of the wording traceable to the page (−15.1 points, interval −25.9 to −4.2), but that did not survive the correction. A table does not seem to get a page chosen, but once chosen, an answer draws more of its wording from it, plausibly because tables hold the specific figures an answer repeats.

## How this compares with other studies

| Source | Design and date | Finding |
|---|---|---|
| This study, 1.1 | 3,096 top-10 pages, compared within the same search, September 2026 | Position first; only a machine-readable date survives correction, and only on informational searches |
| Aggarwal et al. (GEO) | Lab rewrites of sources already given to the engine, 2023 | Visibility up by as much as 40% when the source is in the engine’s context |
| Ahrefs | 1,885 pages that added schema against 4,000 controls, August 2025 to March 2026 | “Adding schema produced no major uplift in citations on any platform”; −4.6% on AI Overviews |
| Seer Interactive | 8,500 keywords, May 2026 | “FAQ + HowTo schema is irrelevant for AIO” |
| SE Ranking (reported by Search Engine Journal) | 216,524 pages, ChatGPT citations, November 2025 | Pages with FAQ schema averaged 3.6 citations against 4.2 without; question-style headings 3.4 against 4.3 |
| Ahrefs | 16.975M cited URLs, July 2025 | AI Overviews and organic results are the most likely to cite older pages |

Version 1.1 brings this study closer to the before-and-after evidence. Ahrefs’ test, the strongest design here, found that adding schema did not raise citations; our schema associations mostly shrink toward zero once pages are compared within the same search and with each other. The version 1.0 finding on question headings, which ran against SE Ranking’s ChatGPT data, largely came from comparing pages across different kinds of searches.

Sources: [Aggarwal et al., GEO](https://arxiv.org/abs/2311.09735); [Ahrefs, schema test](https://ahrefs.com/blog/schema-ai-citations/); [Seer Interactive](https://www.seerinteractive.com/insights/what-it-takes-to-rank-in-googles-ai-overviews-in-2026-is-not-what-you-think); [Search Engine Journal on SE Ranking](https://www.searchenginejournal.com/new-data-top-factors-influencing-chatgpt-citations/561954/); [Ahrefs, freshness](https://ahrefs.com/blog/do-ai-assistants-prefer-to-cite-fresh-content/).

## What this means

This section is interpretation. Because the study is observational, none of these points is a proven cause.

- **Ranking is still the way in.** Among pages that already rank, position explained more of which were cited than all 14 page features together, and most of the variation was explained by neither.
- **There is no markup shortcut.** After comparing like with like, schema types, author signals and FAQ sections showed no reliable association with citation. Adding them is cheap, but the data gives no reason to expect a lift.
- **Match the page to the kind of question.** On informational searches, dated article pages were cited far more often; on commercial searches they were not. The useful question is what shape of page answers this search, not which tags to add.
- **Being cited and being used are different.** Most cited pages were one source among about eight, and for most of them the answer’s wording could not be traced closely to the page. Specific, quotable content such as a table of figures went with more of the answer coming from the page.

## What a causal test would need

The associations above cannot say whether changing a page would change its citations. A test that could would change one thing at a time on real pages, with the rest held constant.

- Pairs of comparable pages, one left as it is and one given a single change (a date, a question heading, a table), assigned at random.
- The same searches run several times per date, and on several dates, because AI Overviews vary from run to run.
- Searches held out from the design, reworded searches, and more than one AI engine, to test whether an effect generalizes.
- All three stages measured, plus a check that the answer states what the page says.
- A version where competing pages are changed too, to tell an absolute gain from one that only holds while competitors stay the same.

We have not run such a test; this study provides the observational baseline for one.

## Methodology

- **Pages:** all top-10 organic results for the 486 US keywords (from our 800-keyword sample) that showed an AI Overview on 26 September 2026, excluding YouTube, Reddit, Facebook, Google, Instagram, TikTok, X, LinkedIn, Quora and Pinterest pages.
- **Cited:** the page’s normalized URL appears among the AI Overview’s references or links on the same results page.
- **Fetching:** direct HTTP with an identifying research user agent; Firecrawl for pages that returned an error or an empty page. Included if the result was HTML of at least 2,000 characters.
- **Features:** JSON-LD types parsed recursively; dates from JSON-LD dateModified, datePublished and uploadDate, article meta times and HTML time elements (newest valid date); question-form heading = an H2 or H3 ending in a question mark; word count on visible text.
- **Stage 1 model:** linear probability model of cited on the 14 features, organic-position dummies and word-count quartile, with search fixed effects, so each page is compared only with pages on the same results page. Features entered together and one at a time.
- **Stage 2:** first source = the first distinct URL linked in the AI Overview text; passage share = share of the answer’s cited passages that link the page.
- **Stage 3:** share of each citing passage’s content words (stop words removed, three or more letters) found in the page’s visible text; absorbed at 80% or more; baseline = the same passages against uncited top-10 pages of the same search.
- **Intent:** DataForSEO keyword-intent labels.
- **Inference:** 95% intervals from 1,000 bootstrap resamples of whole searches (seed 20260928; version 1.0 used seed 20260926); Holm-adjusted p-values across the 14 features; a negative result requires the interval inside ±5 points and an adjusted p above 0.05.
- **Version 1.0 method:** difference between cited and uncited shares within position bands 1 to 3, 4 to 6 and 7 to 10, averaged with band-size weights.
- **Update schedule:** quarterly.

## Limitations

- Correlation, not causation: the study measures how often pages with a trait are cited, not what happens when the trait is added.
- One results page per search on one day; run-to-run and week-to-week stability are not measured, so small estimates may not repeat.
- Retrieval is not observed: pages outside the top 10, and Google’s candidate set, are not seen.
- Absorption is a word-overlap proxy. It does not check that the answer is faithful to the page, and heavy paraphrase lowers it.
- Site authority, topic and content quality are not controlled beyond comparison within the same search and position.
- Intent groups rest on keyword-database labels; the navigational group (31 searches) was too small to analyze separately.
- Declared dates can be changed without changing content.
- Pages that errored, returned a 400+ status, or gave either fetcher too little HTML (14.6%) are excluded; many of these are behind bot protection.

## Data and downloads

- Every page with its features, position and citation status: [s10_pages.csv](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_pages.csv) and [JSON](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_pages.json)
- Every cited page with its prominence and absorption scores: [s10_v11_cited_pages.csv](https://underneath.agency/research-data/ai-overview-cited-pages-study/s10_v11_cited_pages.csv)
- Every statistic on this page, version 1.0 and 1.1: [stats.json](https://underneath.agency/research-data/ai-overview-cited-pages-study/stats.json)
- Machine-readable methodology: [methodology.json](https://underneath.agency/research-data/ai-overview-cited-pages-study/methodology.json)

The data is free to reuse with attribution (CC BY 4.0).

To cite: Underneath. (2026). *What pages cited by AI Overviews have in common: 3,096 pages* (Version 1.1). Underneath Research. https://underneath.agency/research/ai-overview-cited-pages-study

## Frequently asked questions

### Does schema markup help a page get cited in AI Overviews?

Not reliably. Compared with pages on the same search and with other features held constant, no schema type survived a correction for multiple testing, and Organization and BreadcrumbList schema showed no meaningful association at all. A controlled test by Ahrefs also found that adding schema did not raise citations.

### What matters most for being cited in an AI Overview?

Ranking position. 41.7% of pages in positions 1 to 3 were cited, against 20.1% in positions 7 to 10. Among on-page features, only a machine-readable date held up (+7.9 points), and only on informational searches.

### Do question-form headings help?

The evidence is weak. Version 1.0 found a +9.6 point difference, but comparing pages within the same search, with other features held constant, reduced it to +3.5 points with an interval that includes zero.

### Do AI Overviews prefer fresh content?

Having a machine-readable date went with citation (+7.9 points), mainly on informational searches. Being dated within the last 90 days added nothing beyond that (−4.0 points, interval −8.7 to 0.8).

### Does being cited mean the AI Overview used my page?

Not necessarily. Only 21.0% of cited pages were the first source, and in 62.2% of cited pages no passage citing the page had at least 80% of its content words on the page. Pages with tables had more of the answer’s wording traceable to them.

## Related research

- [Do AI Overviews cite the pages that rank?](https://underneath.agency/research/ai-overview-citations-study)
- [How often do AI Overviews appear?](https://underneath.agency/research/ai-overviews-frequency-study)
- [Do websites serve Markdown to AI agents?](https://underneath.agency/research/agent-readable-web-study)
- [AI Mode vs AI Overviews: how different are the sources?](https://underneath.agency/research/ai-mode-vs-ai-overviews-study)

## Related guides

- [What on-page signals are linked to citations in Google AI Overviews and Perplexity?](https://underneath.agency/resources/on-page-signals-linked-to-ai-citations)
- [Can restructuring existing content, without changing what it says, increase AI citations?](https://underneath.agency/resources/does-content-structure-increase-ai-citations)
- [Do FAQ pages help you get cited by AI search engines?](https://underneath.agency/resources/do-faq-pages-help-ai-citations)
- [What does content that AI engines prefer look like?](https://underneath.agency/resources/what-content-do-ai-engines-prefer)
- [What decides whether an AI engine cites my page over a competitor’s?](https://underneath.agency/resources/why-ai-cites-competitor-page-first)

---

This is the Markdown twin of https://underneath.agency/research/ai-overview-cited-pages-study. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
