Guide · AI search

How do we prove that GEO caused a change in sales or conversions?

You need a comparison: pages, products or markets that did not get the GEO work, measured over the same weeks, ideally chosen at random. AI traffic is growing for everyone, so a before-and-after rise mostly measures that growth. The clearest field test so far found a real-looking effect that shrank, and became uncertain, under closer checks.

The short version

  1. On one site, total ChatGPT referrals grew 5.7 times, but pages that got no optimization grew 3.5 times anyway (Watanabe and Nakayashiki, 2026).
  2. After netting out that growth, the optimized pages gained about 1.82 times more ChatGPT referrals, an effect the authors call suggestive, not conclusive.
  3. In a reanalysis of the same data, the estimate ranged from 1.183 to 2.404 times, depending on the assumed start week (Kato and colleagues, 2026).
  4. In 500 simulated runs, a simple before-and-after comparison never captured the true effect when all units shared platform growth (Kato and colleagues, 2026).
  5. A review of GEO research found traffic and conversions to be its weakest evidence; one reported 20% traffic lift lacked the detail to judge it (Martinez, 2026).

Why can’t you just compare sales before and after the GEO work?

Because AI platforms are growing fast, so traffic from them rises whether or not you do anything.

Generative engine optimization, or GEO, is work to make AI engines name and cite you more. Proof matters more as buying moves into AI. A retail survey cited by Wen and colleagues (opens in a new tab) found that 39% of 8,350 shoppers in 21 countries use AI for product discovery and related tasks.

The trouble is the tailwind. Watanabe and Nakayashiki (opens in a new tab) optimized one section of their company’s website, glasp.co, in January 2026 and left the rest alone. Total ChatGPT referrals grew 5.7 times on monthly figures. But the untouched pages grew 3.5 times over the same window. A naive reading would have credited the optimized pages’ 6.1-times growth entirely to the work.

Headline numbers can be even worse. Measured from the lowest day to a single peak day, the site’s ChatGPT traffic rose by a factor of 38.2. The authors themselves call that a fragile basis for any claim.

What did the clearest field test so far find?

A likely but unproven effect. The optimized pages gained roughly 1.8 to 2.3 times relative to untouched pages.

The authors tracked the weekly ratio of optimized to untouched pages’ ChatGPT referrals, before and after the change. Because both groups shared the same site, analytics and platform growth, the ratio cancels much of the common rise. They found a jump of about 1.82 times at the time of the change, and other versions of the analysis all landed between 1.8 and 2.3 times.

Two cautions apply. First, a stricter check found that the months before the change contained random jumps nearly as large, so the authors call the effect suggestive, not conclusive. Second, this is one site, mostly one AI engine, and the authors work for the company that owns it. The optimized and untouched pages were also different kinds of content.

The study also checked for side effects on Google. Google clicks to the optimized pages fell about 25%, close to the roughly 20% fall across the whole site. The authors read that as no measurable harm to Google traffic. Note that these outcomes are referrals and clicks, not sales.

How fragile are estimates like this?

Very. Small, reasonable choices in the analysis can double the estimate or erase it.

Kato and colleagues (opens in a new tab) reran the same public referral data. The rollout happened over several weeks, so the start date is a judgment call. Moving the assumed first week from December 16, 2025 to January 6, 2026 changed the estimated jump from 1.183 times to 2.404 times. For the earlier dates, the plausible range included no effect at all.

A second problem was measurement. In mid-March the site changed how it filtered out bot traffic. Around then, the engagement rate of ChatGPT visitors rose from 0.486 to 0.896, according to the original study. When Kato and colleagues allowed for that change, the estimates fell to between 1.337 and 1.455 times. In one version the estimate was 1.387 times, with a plausible range that included no effect.

What does a proper test of GEO’s business effect need?

A comparison group that shares the same market changes, and ideally random assignment. Kato and colleagues spell out the conditions.

A randomized rollout gives a fair comparison by design: you pick, at random, which products, pages or regions get the GEO work. Without that, you need data from before the change and a comparison group measured at the same time. Both groups must also face the same shifts in the AI engines. In their simulation, across 500 runs where every unit shared platform growth, comparing treated units with themselves before and after never once captured the true effect. Comparing the change in treated units with the change in untreated ones removed the error.

They also show how to connect visibility to sales. How often answers mention you must be multiplied by how many relevant questions are asked, which AI systems people use, and how often readers notice your name. Each piece varies. In their own collection of 2,240 answers, one brand appeared in 33.8% of answers from one OpenAI model and 27.8% from another. The order reversed for English questions. Their method was tested on simulated sales data, not a real company’s.

How strong is the evidence that GEO lifts sales today?

Weak. No study we reviewed shows GEO causing a measured change in sales or conversions.

In a review of the field, Martinez (opens in a new tab) ranks traffic and conversions as the weakest evidence. One industry study reported a 20% production traffic lift against a control, but did not describe group sizes, how units were assigned or the uncertainty. The review found no technique with a proven lasting effect on downstream clicks and conversions. That does not mean GEO does not work. It means the claims have outrun the evidence, so your own test matters.

What should you do about it?

Design the proof before you spend, and make it a fair comparison. A workable plan:

  1. Pick one business outcome (leads, sign-ups or sales from AI referrals) and define it before the work starts.
  2. Split your scope. Choose similar products, pages, categories or regions, and assign some to GEO work and some to wait, at random if you can.
  3. Collect at least several months of history for both groups, so you can see their normal trends.
  4. Track both groups weekly over the same period, and compare the change in one with the change in the other. Server logs of AI bot visits can show which pages in each group the bots request.
  5. Log every other change: analytics settings, bot filtering, site launches, campaigns and price moves. Any of these can fake or hide an effect.
  6. Test the start date. If the result only appears for one chosen week, treat it as unproven.
  7. Discount headline multiples from vendors, agencies or case studies that show no comparison group. The same caution applies to “best of” lists that rank their publisher first.

If you want help setting up a test like this, see our generative engine optimization service.

What does the research not tell us yet?

The research has not yet measured GEO’s effect on sales or revenue for any real business. Open questions:

  • Sales, not referrals. The one field test measured ChatGPT referral visits and Google clicks, not purchases.
  • More than one site. That test covered a single domain and mostly one engine.
  • Which tactic works. Its changes were a bundle, so no single tactic can be credited.
  • Whether people notice. Linking AI mentions to sales needs data on how often readers notice a brand name, and few studies collect it.
  • Random trials. No randomized rollout of GEO with business outcomes has been published.

Frequently asked questions

Can we attribute AI referral traffic growth to our GEO work?

Not without a comparison group. On one site, pages that received no optimization still grew their ChatGPT referrals 3.5 times over the same months.

What is a holdout test for GEO?

It is a test where some comparable products, pages or regions are deliberately left out of the GEO work and measured alongside the rest. Assigning them at random gives the fairest comparison.

How much did GEO increase traffic in the clearest study so far?

About 1.8 to 2.3 times more ChatGPT referrals relative to untouched pages, on one site. A stricter check rated that suggestive, and a reanalysis put it between 1.183 and 2.404 times depending on the start date.

Do AI visibility tools measure sales impact?

Usually not. They measure mentions and citations, which still need to be linked to how many people ask, which engines they use and whether they notice your name.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.