Guide · AI search

What GEO practices are actually supported by research?

The research supports a conservative playbook: publish relevant, complete, verifiable, clearly structured pages that AI systems can reach, then measure whether they are found, cited and quoted correctly. Relevance and ranking position have the strongest evidence. Popular tricks such as keyword stuffing, formatting makeovers and blanket schema have weak or negative evidence, and no technique has been shown to raise visibility durably across AI engines.

The short version

  1. A review of 45 studies by Martinez (opens in a new tab) found no technique with a proven, lasting effect on being found across AI engines; relevance and position were the most reproducible levers.
  2. The famous “up to 40%” figure comes from a lab test where the page was already in front of the AI, in the original GEO paper (opens in a new tab).
  3. In one benchmark, only three of 54 combinations of rewrite method and subject area clearly helped.
  4. Pages with numbers, definitions and comparisons shaped AI answers more, in an analysis of 18,151 cited pages (opens in a new tab); question-and-answer formatting alone did not.
  5. Evidence that GEO raises traffic or sales is the weakest of all: one suggestive field study and a few industry claims.

What does the research actually support?

It supports good, findable information pages, measured carefully, not tricks. That is the conclusion of the most complete review of the field so far.

Olivier Martinez reviewed 45 studies published between November 2023 and July 2026. His summary: “produce a relevant, comprehensive, verifiable, clearly structured, and technically retrievable page; then measure retrieval, citation, and fidelity separately.” In plain words, the page must be found, then cited, then represented accurately, and each step needs its own check. For plain definitions of terms like GEO and citations, see our AI search glossary.

The evidence is strongest for one narrow claim. Once a page is already among the sources an AI system has gathered, changing the page can change how it is cited or used. Whether those changes help a page get gathered in the first place is far less proven.

Which practices have the strongest evidence?

Relevance to the question and position among sources have the strongest support. Specific, checkable evidence comes next.

PracticeStrength of evidenceWhat it means for you
Answer the exact question askedStrongWrite for real buyer questions
Rank well enough to be gatheredStrongSearch work still matters
Verifiable figures, definitions, comparisonsModerate to strongAdd real, sourced facts
Prices, dates, recencyModerateHelps on time-sensitive and buying questions
Page structureModerate, mixedTest it; do not assume it helps
Authoritative toneWeak and unstableConfidence is not evidence
Keyword stuffingNone or negativeAvoid
Formatting alone, fixed recipesPoorAvoid paying for makeovers

The grades come from the Martinez review. A controlled test by Vishwakarma and colleagues (opens in a new tab) across 252,000 trials backs the top rows: topic match and position decided which page was cited first.

The content analysis by Zhang and colleagues points the same way for evidence. By the authors’ measure of influence, cited pages containing numbers or statistics shaped answers 61.55% more on average than pages without. Pages in question-and-answer format did slightly worse, by 5.74%.

Is the “40% more visibility” claim true?

Only inside the lab setup that produced it. The figure does not mean 40% more citations, traffic or customers.

The original GEO study reported that rewrites “can boost visibility by up to 40%”. Martinez traces this to one score rising from 19.3 to 27.2 when quotations were added. In that test, five pages were already placed in front of the AI, so the result says nothing about being found.

The study’s live check on Perplexity used 200 examples, uploaded as files rather than found on the web. Martinez lists “GEO increases visibility by 40%” as a claim rejected in general form.

Why do results vary so much between studies?

AI engines differ from each other and change constantly, so one snapshot rarely generalizes. Even repeated runs of the same question give different sources.

Martinez cites a two-month study in the US and Germany where only 18% of the pages in AI Overviews stayed the same, against 45% for regular Google results. Repeated runs set up to be as consistent as possible still changed 9–28% of decisions. In another study, 57.8% of ChatGPT runs did not search the web at all.

Engines also disagree with each other. In our comparison of Google’s two AI surfaces, AI Mode and AI Overviews shared a mean 13.9% of cited URLs on the same searches. A result measured on one engine, on one day, should not be treated as a rule.

Does GEO lead to more traffic or sales?

The evidence here is the weakest in the field. One field study suggests a gain, but it is not proof.

In that study, ChatGPT referrals to treated pages rose by a factor of 5.7. But untreated pages on the same site had already risen 3.5 times as the platform grew. A careful analysis put the extra effect at 1.82 times, which Martinez calls “suggestive rather than causally established”. One industry study reported a 20% traffic lift, without enough detail to judge it.

Being cited is not the same as being represented well, either. An early audit found only 51.5% of sentences in AI search answers were fully supported by their citations. Our article on AI Mode and traffic covers what the shift to AI answers may do to clicks.

What should you do about it?

Follow the conservative playbook and measure each stage separately. In practice:

  1. Map the real questions buyers ask, and make sure one page answers each directly.
  2. Back claims with specific, sourced figures, definitions and comparisons. Never invent statistics or reviews.
  3. Use clear headings and tables, but judge them by results, not by assumption.
  4. Keep pages reachable. In our crawler study, 7.4% of top sites blocked ChatGPT’s search crawler.
  5. Keep investing in search ranking, since being gathered comes before being cited.
  6. Measure repeatedly: run each question several times, on several dates and engines, and track found, cited and quoted correctly as separate numbers.
  7. Avoid hidden instructions to AI systems, fake authority and undisclosed promotion; Martinez’s tests for honest optimization rule these out.

For help building this kind of measured program, see our generative engine optimization service.

What does the research not tell us yet?

The research has not shown that any GEO practice durably lifts visibility, traffic or revenue across engines.

  • Being found. Very few studies test whether a change helps a real page get retrieved by a live engine.
  • Competition. Gains may shrink as everyone adopts the same tactics; C-SEO Bench found the problem behaves like a zero-sum game.
  • Business outcomes. Clicks, leads and sales are almost never measured with a proper control group.
  • Durability. Engines change often, and most studies are single snapshots.
  • Non-English markets. Most studies use English, anonymous accounts and few locations.

Frequently asked questions

Is GEO proven to work?

Partly. Changing a page already gathered by an AI system can change how it is cited. But no reviewed technique has shown a stable, cross-engine effect on being found or on traffic.

Does keyword stuffing work for AI search?

No. In the original GEO study, keyword stuffing performed 10% worse than the unedited page on Perplexity, and later reviews rate it null or negative.

Is GEO different from SEO?

It builds on SEO rather than replacing it. The strongest levers are relevance and position among sources, and both depend on being found by search first.

How should I measure GEO results?

Measure being found, being cited and being quoted correctly as separate numbers, over repeated runs. In one study, 57.8% of ChatGPT runs did not search the web at all, so a single check can mislead.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.