Guide · AI search

How do we know whether a GEO campaign actually improved our AI citations?

You know only if the gain beats the normal swing in AI answers, lasts for weeks, and is not matched where you changed nothing. A single before-and-after snapshot cannot tell you. In one study’s worked example, a rise in citation share from 8% to 11% on OpenAI’s search model would sit within ordinary noise.

The short version

  1. On OpenAI’s search model, a rise in citation share from 8% to 11% would fall within the usual margin of error, so it would prove nothing (Sielinski, 2026).
  2. Across three engines and three topics, websites whose citation shares differed by less than 5 to 7 points usually could not be told apart (Sielinski, 2026).
  3. With nothing changed, the top cited website on one topic swung by a factor of nearly 2 in nine days (Sielinski, 2026).
  4. In one tracking platform’s data, brands that did nothing drifted down about 1.34 points per run on ChatGPT, so “before” is not a fixed line (Kumar, 2026).
  5. A review of 45 studies found no technique with a proven lasting effect on organic AI visibility across engines (Martinez, 2026).

Can a before-and-after comparison prove a GEO campaign worked?

Not on its own. AI answers change between runs, so two snapshots mostly measure noise.

Generative engine optimization, or GEO, means changing your content and presence so AI engines name and cite you more. The trouble is that the same question asked twice gives different sources. Sielinski (opens in a new tab) sent the same queries to Gemini, Perplexity and OpenAI’s search model every day for nine days. He also checked that the cited pages themselves had not changed. The swings came mostly from the engines, not from edits to the web.

His conclusion is blunt. If your share of OpenAI’s citations rises from 8% to 11% after a campaign, the gain cannot be attributed to the campaign with confidence. The 3-point rise sits inside the typical margin of error for a website on that engine. Only larger effects, or effects confirmed by repeated measurement before and after, stand out. No real campaign was tested here, the topics were consumer products, and the author works for a company that sells AI visibility measurement.

How big does a change need to be before it counts?

Bigger than most people expect. Gaps of a few points between two measurements are often within the noise.

Sielinski gives an example from running gear. With 200 queries, Tom’s Guide had about 9.5% of citations on OpenAI’s search model and Runner’s World about 6.0%. That looks like a clear lead. But the plausible ranges for the two sites overlapped almost entirely. Across his platforms and topics, sites that seemed to differ by less than 5 to 7 points usually could not be separated.

The same problem applies to brand mentions. In our consistency study, a brand named in 3 of 5 runs could truly appear anywhere from 23.1% to 88.2% of the time. A move from 2 of 5 to 3 of 5 tells you almost nothing. The same caution applies to a one-time AI visibility report.

Can a change in rankings prove progress?

Rarely. A cited website’s rank is itself uncertain, even among the most-cited sites.

In a second dataset of 10 topics, Sielinski (opens in a new tab) measured how much each site’s rank could plausibly vary. Even among the ten most-cited websites, the typical range spanned 5.0 rank positions, and 18.7% of them had a range wider than ten positions. A site moving from rank 10 to rank 20 between two periods may have lost ground, or may simply be bouncing within that noise.

Why does timing matter so much?

Because AI answers drift on their own, within hours and over weeks. A change you see may have happened without you.

Several studies show movement with no campaign at all:

  • Over days. On OpenAI’s search model, the most-cited website for multivitamins swung by a factor of nearly 2 within nine days (Sielinski (opens in a new tab)).
  • Day to day. In four Swiss industries, about 65% of cited sources changed from one day to the next. Even a 14-day window was only enough to show direction, not to compare brands finely (Schulte and colleagues (opens in a new tab)).
  • Within hours. In our consistency study, a ChatGPT answer to the same question 4.4 hours later overlapped less with the earlier answers than they did with each other.
  • Slow decline. On one tracking platform, brands that took no action lost about 1.34 points of visibility per run on ChatGPT (Kumar (opens in a new tab)). The author co-founded the platform, and some of the drift may come from prompt changes.

The last point cuts both ways. If visibility was drifting down, holding steady after a campaign may itself be a gain. A simple before-and-after would miss it.

What does a credible test of a GEO campaign look like?

One that compares changed pages or prompts with similar ones left alone, measured repeatedly over the same weeks.

Martinez (opens in a new tab) reviewed 45 studies and recommends a design any team can borrow. Fix the measure and the threshold for success before you start. Measure an untreated baseline. Ask each question several times and in several wordings, across named engines and dates. Where you can, assign the change at random to some pages or topics and not others.

A control group matters because AI traffic is growing everywhere. In a field study by Watanabe and Nakayashiki (opens in a new tab), pages that received no changes still grew their ChatGPT referrals 3.5 times over the same period. Without that comparison, the whole rise would have been credited to the optimization. We list the other rival explanations in whether GEO work raised your visibility.

It also helps to watch the right unit. On Kumar’s platform, 77.5% of brand, prompt and engine combinations were either always or never mentioned across runs. A prompt that moves from never mentioning you to always mentioning you is a stronger signal than a small shift in an average.

What should you do about it?

Treat every GEO campaign as an experiment, and design the measurement before the work starts. In practice:

  1. Write down the success measure and the minimum gain that would count, before any change goes live.
  2. Collect a baseline over several weeks, with each prompt asked several times, so you know the normal swing.
  3. Hold back a control. Leave comparable pages, products or prompt topics untouched, and track them alongside the ones you change.
  4. Measure after the change for as long as before, using two-to-four-week averages rather than daily readings.
  5. Report ranges, not points. “Up from 8% to 11%” means little without the margin of error around each figure.
  6. Check each engine separately. A gain on Perplexity says nothing about ChatGPT. Our guide to measuring share of citations covers the setup.
  7. Be suspicious of any result that only one snapshot supports, including a vendor’s or an agency’s.

If you want help setting up a test like this, see our generative engine optimization service.

What does the research not tell us yet?

The research has not shown which GEO techniques reliably raise AI citations in live engines over time. Key gaps:

  • Real campaign tests. The 8%-to-11% example is a worked illustration, not a tested campaign.
  • Lasting effects. Martinez found no technique with a stable, long-term effect on organic visibility across engines.
  • The right threshold. How big a gain must be depends on engine, topic and sample size, and no general table exists yet.
  • Competitor effects. If rivals optimize too, your share can fall even when your own work succeeds. Few studies measure this.
  • Independent replication. Several of the key studies come from companies that sell measurement or from the site being studied.

Frequently asked questions

How long should a GEO test run before we judge it?

At least several weeks before and after the change. In the Swiss study, even a 14-day window only showed direction, and the authors recommend rolling averages over two to four weeks.

Is a 3-point increase in AI citation share meaningful?

Often not. On OpenAI’s search model, a rise from 8% to 11% fell within the typical margin of error for a single website.

What is a control group in GEO?

It is a set of pages, products or prompts you deliberately leave unchanged and measure alongside the ones you optimize. In one field study, unchanged pages still grew their ChatGPT referrals 3.5 times, which shows why the comparison matters.

Can a vendor dashboard prove our GEO worked?

Only if it reports repeated runs, ranges around each figure and an untreated comparison. Rankings alone are weak evidence: even top-ten websites had rank ranges spanning 5.0 positions.

Sources

Free strategy call

Some questions are easier to answer about your own business.

Bring the one that matters most. On a free 30-minute call we’ll take a first look at it and send you a short written read afterward.