The short version
- Dates, clean HTML structure and structured data were the signals most linked to citation in an audit of 1,100 B2B software pages (opens in a new tab) cited by Brave, Google and Perplexity.
- In our study of 3,096 ranking pages, a machine-readable date was the only page feature still linked to AI Overview citation after a fair comparison: +7.9 points.
- That date link was confined to informational searches: +17.8 points there, against +1.0 on commercial and transactional searches.
- Perplexity cited much lower-scoring pages than Google: an average quality score of 0.300, against 0.687 for Google AI Overviews, in the same audit.
Which on-page signals are most linked to AI citations?
Dates and freshness metadata, semantic HTML and structured data showed the strongest links. That is the main finding of the GEO-16 audit by Kumar and Palkhouski, at UC Berkeley and Wrodium Research.
They ran 70 prompts aimed at B2B software buyers and collected 1,702 citations from Brave, Google AI Overviews and Perplexity. AI Overviews are the AI summaries at the top of Google’s results. Other terms used here are defined in our plain-language AI search glossary. The team then scored 1,100 unique cited pages on 16 quality “pillars”, from metadata to readability.
Three pillars stood out.
On a scale where 0 means no link and 1 a perfect one, Metadata and Freshness scored 0.68, Semantic HTML 0.65 and Structured Data 0.63. Semantic HTML means one main title and a logical order of headings, so a machine can tell what each part of the page is. Evidence and citations, authority and trust, and internal linking came next.
How strong is that evidence?
It is suggestive, not proof, because the audit compared pages as they were and changed nothing. The authors say so themselves.
They report that pages scoring at least 0.70 overall, with at least 12 of the 16 pillars rated well, reached a 78% cross-engine citation rate. Pages cited by more than one engine (134 URLs) scored 71% higher on quality than pages cited by one engine only.
The limits matter. The pages were English B2B software content, collected at a single point in time.
Brand reputation and backlinks, links from other sites, were not controlled. The authors note they “may influence both scores and citations.” They list schema experiments as future work.
Does schema markup get pages cited in AI Overviews?
The best evidence says schema alone does not. Once pages are compared with competitors on the same search, most schema links shrink toward zero.
In our study of 3,096 top-10 pages on 486 US searches, Organization schema added +0.5 points to the chance of citation and BreadcrumbList schema +1.3, both meaningless. FAQPage schema showed +5.9 points but did not survive a correction for testing 14 features at once.
The strongest design we know of points the same way. Our study summarizes a before-and-after test by the SEO tool maker Ahrefs, of 1,885 pages that added schema against 4,000 controls.
That test found adding schema “produced no major uplift in citations on any platform”. Schema is also common among cited pages anyway. In a study of 615 ChatGPT health citations (opens in a new tab), 74.3% of cited sources used schema markup.
Do dates and freshness help pages get cited?
A machine-readable date helps on informational searches, but a recent date does not add much on Google. On Google, being dated seems to matter more than being new.
In our study, a machine-readable date was the only one of 14 features still clear after the correction, at +7.9 points. On informational searches it rose to +17.8 points. Pages dated within the last 90 days were, if anything, 4.0 points less likely to be cited than other dated pages.
Our freshness study adds a warning about what dates mean. Among pages in Google’s top 10 carrying a recent date, 66.4% were older pages with a new modified date. A lab test by Vishwakarma and colleagues (opens in a new tab) did find a strong pull toward recent dates, but it compared content dated 2026 versus 2019 in a simulation.
| Signal on Google AI Overviews | Linked to citation? | Source |
|---|---|---|
| Machine-readable date | Yes, +7.9 points, mainly informational searches | Our study of 3,096 pages |
| Date within 90 days | No (−4.0 points) | Our study of 3,096 pages |
| Organization or BreadcrumbList schema | No (+0.5 and +1.3) | Our study of 3,096 pages |
| Question-style headings | Not distinguishable from zero (+3.5) | Our study of 3,096 pages |
| HTML table | No for citation (+0.6), yes for wording used | Our study of 3,096 pages |
Is Perplexity different from Google?
Yes: Perplexity cites more sources, and the pages it cites score lower on quality audits. It also cites newer pages than Google ranks.
In the GEO-16 audit, Perplexity’s cited pages averaged a quality score of 0.300, against 0.687 for Google AI Overviews and 0.727 for Brave. In a study of 602 prompts (opens in a new tab), Perplexity cited 16.35 sources per prompt on average, against 12.06 for Google. Each of its cited pages shaped less of the answer than ChatGPT’s sources did.
In our freshness study, 19.2% of Perplexity’s dated citations were pages published in the last 90 days, against 9.0% of Google’s top 10 for the same questions. Access rules also work differently. Our crawler study found 40.4% of the pages Perplexity cited on top sites were closed to its crawler by the site’s own robots.txt file.
Do headings, tables and page length matter?
Not much for being cited, though tables seem to help once a page is cited. Most layout signals vanish when pages are compared fairly.
Question-style headings showed +3.5 points in our study, which is not distinguishable from zero. Word count did not matter once position and search were held constant. Tests that changed layout alone are weighed in whether restructuring content lifts AI citations.
But cited pages with an HTML table were 14.7 points more likely to have the answer’s wording traced back to them. Tables likely hold the specific figures an answer repeats.
Ranking dwarfed all of this. Of top-10 pages, 41.7% ranking 1 to 3 were cited, against 20.1% at positions 7 to 10. Most of the variation, 91.4%, was explained by nothing we measured.
What should you do about it?
Treat these signals as basic hygiene, and put most effort into ranking and answering the question well. Concretely:
- Show a visible date and a matching machine-readable date on informational content, and only update it when the content changes.
- Keep one main title and a logical heading order on every important page.
- Keep structured data valid and matching the visible page, but do not expect it to lift citations alone.
- Put key figures in simple tables, so answers that cite you can quote you accurately.
- Check your robots.txt deliberately for each AI crawler instead of relying on defaults.
- Measure Perplexity and Google separately, because they reward different things.
If you want help prioritizing these across a large site, see our generative engine optimization service.
What does the research not tell us yet?
No published study has changed these signals on real pages and measured citations before and after across engines.
- Cause and effect. The GEO-16 audit and our study both compare existing pages. Better-run sites may simply have better pages and better rankings.
- Other sectors and languages. GEO-16 covered English B2B software pages; our study covered US searches on one day.
- Brand strength. Backlinks and reputation were not controlled in the audit, and the GEO-16 scores were produced by the authors’ own framework.
- Stability. Both studies are snapshots, and engines change often.
- Perplexity in depth. Fewer studies cover Perplexity than Google, and none tested schema there directly.
Frequently asked questions
Does schema markup help with Google AI Overviews?
Not on its own, according to the best current evidence. In our study of 3,096 pages, most schema types showed no reliable link once pages were compared on the same search, and a before-and-after test of 1,885 pages found no major uplift.
Should I add dates to my pages for AI search?
Yes, on informational content. A machine-readable date was the one feature linked to AI Overview citation in our study, worth +17.8 points on informational searches, but a newer date added nothing extra.
Why does Perplexity cite different pages than Google?
Perplexity casts a wider net, and its cited pages score lower on quality audits. It cited 16.35 sources per prompt in one study and cited pages with an average quality score of 0.300, far below Google’s 0.687.
What is semantic HTML?
It is page code that labels each part by its role: one main title, ordered subheadings, lists and tables. It was one of the three signals most linked to AI citation in the GEO-16 audit.
Do author bios help AI citations?
The evidence does not show it. In our Google study, an author signal showed no reliable link, and 64.7% of ChatGPT-cited health sources had no author attribution at all.
Sources
- Kumar and Palkhouski (2025), AI Answer Engine Citation Behavior An Empirical Analysis of the GEO16 Framework (opens in a new tab), arXiv:2509.10762.
- Vishwakarma, Kumar and Jamidar (2026), What Gets Cited: Competitive GEO in AI Answer Engines (opens in a new tab), arXiv:2605.25517.
- Zhang, He and Yao (2026), From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms (opens in a new tab), arXiv:2604.25707.
- Jacques and colleagues (2026), Authority Signals in AI Cited Health Sources: A Framework for Evaluating Source Credibility in ChatGPT Responses (opens in a new tab), arXiv:2601.17109.
- Underneath (2026), What pages cited by AI Overviews have in common: 3,096 pages
- Underneath (2026), How fresh are the pages AI engines cite?
- Underneath (2026), Which AI crawlers do top websites block?