The short version
- A review of 45 studies by Martinez (opens in a new tab) found no technique with a proven, lasting effect on being found across AI engines; relevance and position were the most reproducible levers.
- The famous “up to 40%” figure comes from a lab test where the page was already in front of the AI, in the original GEO paper (opens in a new tab).
- In one benchmark, only three of 54 combinations of rewrite method and subject area clearly helped.
- Pages with numbers, definitions and comparisons shaped AI answers more, in an analysis of 18,151 cited pages (opens in a new tab); question-and-answer formatting alone did not.
- Evidence that GEO raises traffic or sales is the weakest of all: one suggestive field study and a few industry claims.
What does the research actually support?
It supports good, findable information pages, measured carefully, not tricks. That is the conclusion of the most complete review of the field so far.
Olivier Martinez reviewed 45 studies published between November 2023 and July 2026. His summary: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.” In plain words, the page must be found, then cited, then represented accurately, and each step needs its own check. For plain definitions of terms like GEO and citations, see our AI search glossary.
The evidence is strongest for one narrow claim. Once a page is already among the sources an AI system has gathered, changing the page can change how it is cited or used. Whether those changes help a page get gathered in the first place is far less proven.
Which practices have the strongest evidence?
Relevance to the question and position among sources have the strongest support. Specific, checkable evidence comes next.
| Practice | Strength of evidence | What it means for you |
|---|---|---|
| Answer the exact question asked | Strong | Write for real buyer questions |
| Rank well enough to be gathered | Strong | Search work still matters |
| Verifiable figures, definitions, comparisons | Moderate to strong | Add real, sourced facts |
| Prices, dates, recency | Moderate | Helps on time-sensitive and buying questions |
| Page structure | Moderate, mixed | Test it; do not assume it helps |
| Authoritative tone | Weak and unstable | Confidence is not evidence |
| Keyword stuffing | None or negative | Avoid |
| Formatting alone, fixed recipes | Poor | Avoid paying for makeovers |
The grades come from the Martinez review. A controlled test by Vishwakarma and colleagues (opens in a new tab) across 252,000 trials backs the top rows: topic match and position decided which page was cited first.
The content analysis by Zhang and colleagues points the same way for evidence. By the authors’ measure of influence, cited pages containing numbers or statistics shaped answers 61.55% more on average than pages without. Pages in question-and-answer format did slightly worse, by 5.74%.
Is the “40% more visibility” claim true?
Only inside the lab setup that produced it. The figure does not mean 40% more citations, traffic or customers.
The original GEO study reported that rewrites “can boost visibility by up to 40%”. Martinez traces this to one score rising from 19.3 to 27.2 when quotations were added. In that test, five pages were already placed in front of the AI, so the result says nothing about being found.
The study’s live check on Perplexity used 200 examples, uploaded as files rather than found on the web. Martinez lists “GEO increases visibility by 40%” as a claim rejected in general form.
Which popular tactics do not hold up?
Keyword stuffing, one-size-fits-all rewrites and markup shortcuts have weak or negative evidence. Some rewrites even hurt.
- Keyword stuffing. On Perplexity, it performed 10% worse than the original page in the GEO study. Our guide on whether keyword stuffing still works covers the later tests.
- Generic rewrite recipes. In the C-SEO Bench benchmark, only three of 54 method and subject combinations were clearly positive, and some rewrites lowered a page’s rank.
- Rewrites that hurt being found. In an end-to-end test, rewriting only the body text cut a page’s presence among the top 10 candidates by 16%.
- Schema as a shortcut. In our study of AI Overview citations, Organization schema added +0.5 points to the chance of being cited, which is no meaningful difference.
- llms.txt as a cure-all. Only 11.5% of top sites publish a valid file, in our llms.txt study, and no one we found has shown it changes citations.
Why do results vary so much between studies?
AI engines differ from each other and change constantly, so one snapshot rarely generalizes. Even repeated runs of the same question give different sources.
Martinez cites a two-month study in the US and Germany where only 18% of the pages in AI Overviews stayed the same, against 45% for regular Google results. Repeated runs set up to be as consistent as possible still changed 9–28% of decisions. In another study, 57.8% of ChatGPT runs did not search the web at all.
Engines also disagree with each other. In our comparison of Google’s two AI surfaces, AI Mode and AI Overviews shared a mean 13.9% of cited URLs on the same searches. A result measured on one engine, on one day, should not be treated as a rule.
Does GEO lead to more traffic or sales?
The evidence here is the weakest in the field. One field study suggests a gain, but it is not proof.
In that study, ChatGPT referrals to treated pages rose by a factor of 5.7. But untreated pages on the same site had already risen 3.5 times as the platform grew. A careful analysis put the extra effect at 1.82 times, which Martinez calls “suggestive rather than causally established”. One industry study reported a 20% traffic lift, without enough detail to judge it.
Being cited is not the same as being represented well, either. An early audit found only 51.5% of sentences in AI search answers were fully supported by their citations. Our article on AI Mode and traffic covers what the shift to AI answers may do to clicks.
What should you do about it?
Follow the conservative playbook and measure each stage separately. In practice:
- Map the real questions buyers ask, and make sure one page answers each directly.
- Back claims with specific, sourced figures, definitions and comparisons. Never invent statistics or reviews.
- Use clear headings and tables, but judge them by results, not by assumption.
- Keep pages reachable. In our crawler study, 7.4% of top sites blocked ChatGPT’s search crawler.
- Keep investing in search ranking, since being gathered comes before being cited.
- Measure repeatedly: run each question several times, on several dates and engines, and track found, cited and quoted correctly as separate numbers.
- Avoid hidden instructions to AI systems, fake authority and undisclosed promotion; Martinez’s tests for honest optimization rule these out.
For help building this kind of measured program, see our generative engine optimization service.
What does the research not tell us yet?
The research has not shown that any GEO practice durably lifts visibility, traffic or revenue across engines.
- Being found. Very few studies test whether a change helps a real page get retrieved by a live engine.
- Competition. Gains may shrink as everyone adopts the same tactics; C-SEO Bench found the problem behaves like a zero-sum game.
- Business outcomes. Clicks, leads and sales are almost never measured with a proper control group.
- Durability. Engines change often, and most studies are single snapshots.
- Non-English markets. Most studies use English, anonymous accounts and few locations.
Frequently asked questions
Is GEO proven to work?
Partly. Changing a page already gathered by an AI system can change how it is cited. But no reviewed technique has shown a stable, cross-engine effect on being found or on traffic.
Does keyword stuffing work for AI search?
No. In the original GEO study, keyword stuffing performed 10% worse than the unedited page on Perplexity, and later reviews rate it null or negative.
Is GEO different from SEO?
It builds on SEO rather than replacing it. The strongest levers are relevance and position among sources, and both depend on being found by search first.
How should I measure GEO results?
Measure being found, being cited and being quoted correctly as separate numbers, over repeated runs. In one study, 57.8% of ChatGPT runs did not search the web at all, so a single check can mislead.
Sources
- Martinez (2026), Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026) (opens in a new tab), arXiv:2607.14035.
- Aggarwal and colleagues (2023), GEO: Generative Engine Optimization (opens in a new tab), arXiv:2311.09735.
- Puerto and colleagues (2025), C-SEO Bench: Does Conversational SEO Work? (opens in a new tab), arXiv:2506.11097.
- Vishwakarma, Kumar and Jamidar (2026), What Gets Cited: Competitive GEO in AI Answer Engines (opens in a new tab), arXiv:2605.25517.
- Zhang, He and Yao (2026), From Citation Selection to Citation Absorption: A Measurement Framework for Generative Engine Optimization Across AI Search Platforms (opens in a new tab), arXiv:2604.25707.
- Underneath (2026), What pages cited by AI Overviews have in common: 3,096 pages
- Underneath (2026), How many websites have an llms.txt file? 2026 adoption data
- Underneath (2026), AI Mode vs AI Overviews: how different are the sources?
- Underneath (2026), Which AI crawlers do top websites block? 10,000 sites, 2026