---
title: "Our AI visibility went up: was it because of our GEO work?"
description: "Maybe. A rise can also come from chance, AI platform growth, engine updates, competitors or a changed prompt set, so you need a baseline and a control."
canonical: "https://underneath.agency/resources/did-geo-work-raise-ai-visibility"
published: 2026-10-07
updated: 2026-10-08
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
Guide · AI search

# Our AI visibility went up: was it because of our GEO work?

You cannot tell from the rise alone. Your own work is only one of at least four things that move an AI visibility score, alongside competitors’ changes, updates to the AI engines and changes to the set of questions being tracked. Ordinary run-to-run chance can produce swings of several points on top of that. Before you credit or blame a GEO program, you need a baseline measured the same way and a comparison group that did not get the work.

## The short version

1. A survey of GEO measurement lists at least four causes of a score change besides wording and scoring: your work, competitors’ work, engine drift and a change in the question set ([Martinez, 2026](https://arxiv.org/abs/2609.06811)).
2. On one website, ChatGPT referrals to optimized pages grew 6.1 times, but untouched pages on the same site grew 3.5 times; the part plausibly due to the work was about 1.8 to 2.3 times ([Watanabe and Nakayashiki, 2026](https://arxiv.org/abs/2606.04362)).
3. In one study, a rise in a site’s ChatGPT search citation share from 8% to 11% was within normal noise and could not be credited to any change ([Sielinski, 2026](https://arxiv.org/abs/2603.08924)).
4. A review of 45 GEO studies found no technique with a proven, lasting effect across AI platforms on being found organically ([Martinez, 2026](https://arxiv.org/abs/2607.14035)).

## What else could have raised our AI visibility?

At least four things besides your own work. [Martinez](https://arxiv.org/abs/2609.06811), in a critical survey of how GEO visibility is measured, lists them: your own intervention, competitors’ interventions, changes in the AI engine itself, and a change in the mix of questions being tracked. Changes in how questions are worded or answers are scored come on top.

The survey is explicit about the limit. A publisher sees a change in its score over time, not the effect of its own work. Comparing two dates, without recording what competitors did and which engine version answered, mixes all of these causes together. Even fixing the prompts does not remove the effect of competitors, because their pages change what the engine finds.

The last cause is easy to miss. If your tracking tool added questions, dropped some or reweighted them, the score can move even if no AI answer changed. Ask your vendor for the question list behind both readings.

## How much of a rise can be pure noise?

More than most dashboards admit: several percentage points, sometimes more. [Sielinski](https://arxiv.org/abs/2603.08924) asked three AI search engines the same 200 questions per topic every day for nine days. On ChatGPT search, a rise in a site’s citation share from 8% to 11% fell within the normal range of noise and could not be credited to an intervention.

Single readings can mislead badly. On one day, Gemini gave nationalgeographic.com a citation share of 0.032 for bird feeder questions; across the nine days its average was 0.005. Yahoo’s share for multivitamin questions on ChatGPT search ranged from 0.092 to 0.160 within the window, nearly a twofold swing for the top site. The author checked that the cited pages themselves had mostly not changed. That is why [a one-time visibility report](https://underneath.agency/resources/one-time-ai-visibility-report) needs repeated runs behind it.

Time alone moves answers too. In our [consistency study](https://underneath.agency/research/ai-recommendation-consistency-study), the same ChatGPT question asked about 4.4 hours later overlapped with the earlier runs by 0.477, against 0.588 between runs minutes apart. Sielinski is affiliated with the company IQRush; the consistency study is our own.

## How much of the growth is just AI platforms growing?

Possibly most of it, if you measure AI referral traffic. [Watanabe and Nakayashiki](https://arxiv.org/abs/2606.04362) of Glasp studied their own site after a bundle of changes to one section in January 2026. The rest of the site was left alone and served as a comparison group.

Total ChatGPT referrals to the site grew 5.7 times on monthly figures. The optimized pages grew 6.1 times, but the untouched pages grew 3.5 times with no work at all. Measured from the lowest day to the single best day, the rise looked like 38.2 times, a figure the authors call a fragile basis for any claim.

Their best estimate of the effect of the work was a jump of about 1.8 to 2.3 times. Even that is only suggestive: a stricter check found jumps of similar size at made-up dates before the change. The authors suspect many published success stories are inflated the same way. This is one site, one engine and a bundle of changes, written up by the company that made them. For the traffic side in more detail, see [crediting ChatGPT referral growth](https://underneath.agency/resources/chatgpt-referral-growth-and-geo).

## Can competitors or engine updates move your score?

Yes, and without any change on your side. Citation share is a zero-sum measure: when one source gains share, another loses it. A [review of 45 GEO studies](https://arxiv.org/abs/2607.14035) notes that tested gains can shrink as more competitors adopt the same tactics, and that a before-and-after comparison without a matched baseline confuses the treatment with chance.

Engines also drift. Ranqo, which sells AI visibility tracking, followed brands that did little or nothing to change their visibility. On ChatGPT, their visibility for questions that did not name them slipped by about 1.34% per tracking run ([Kumar, 2026](https://arxiv.org/abs/2606.20065)). The author notes some of that may reflect changes to the question set, not the engine.

Short-term turnover is large. In a test of 15 commercial prompts, the exact pages an engine cited for the same prompt changed by 67.0% from one day to the next, on average across three engines ([Tannenbaum, 2026](https://arxiv.org/abs/2609.22655)). The author, who is affiliated with Aiso Boost Ltd., warns the figure may partly reflect changes in his own collection system.

## How can you tell whether your GEO work caused the change?

Compare against something that did not get the work, measured the same way at the same time. The studies above point to a few designs that are practical for a marketing team.

| Design | What it rules out | Source |
|---|---|---|
| Untouched pages or products on the same site as a comparison group | Platform growth and broad engine changes | Watanabe and Nakayashiki |
| Randomly holding back some pages from the changes | Most other causes, if the groups are large enough | Watanabe and Nakayashiki (proposed, not run) |
| A frozen question set measured before and after | Changes in the question mix | Martinez |
| Many runs before and after, reported with ranges | Run-to-run chance | Sielinski |
| Long enough windows | Day-to-day swings | Schulte and colleagues |

On window length, a Swiss study of four AI engines found per-brand estimates needed about 24 days of daily runs to become reasonably stable ([Schulte, Bleeker and Kaufmann, 2026](https://arxiv.org/abs/2604.07585)). Our guide to [testing whether GEO improved citations](https://underneath.agency/resources/did-geo-improve-ai-citations) shows how to set the threshold for a real gain.

## What should you do about it?

Set up the comparison before you start the work, not after the number moves. Steps:

1. Freeze your question list, assistants, locations and run counts before the program starts, and keep a version history of any change.
2. Measure a baseline of at least several weeks, with repeated runs, so you know your normal range.
3. Keep a comparison group: pages, products or topics you deliberately leave alone. Where you can, choose them at random.
4. Track the competitors named alongside you, and note known engine updates on your timeline.
5. Report the change as a range, and give the effect on treated items minus the change in the comparison group.
6. For traffic, compare AI referrals to treated pages with untreated pages on the same site, as in the Glasp study.

If you want a program measured this way from the start, see our [generative engine optimization service](https://underneath.agency/services/generative-engine-optimization).

## What does the research not tell us yet?

The research shows how a rise can mislead, but not yet how much GEO work reliably delivers. Gaps:

- The review of 45 studies found that content changes can affect answers once a page is already retrieved. It found no technique with a proven, lasting, cross-platform effect on being found or on clicks and sales.
- The best traffic evidence comes from one site, one engine and one bundle of changes, reported by the site’s owner.
- No study has yet split an observed change into its four causes with real data. The survey that names them offers a framework, not measurements.
- Most noise figures come from consumer product topics over days or weeks. Long-run engine drift through major model updates is barely measured.

## Frequently asked questions

### How do I know if GEO improved our AI visibility?

Compare the change on the pages or topics you worked on with a similar group you left alone, measured with the same questions and run counts. Without that comparison, platform growth and noise can look like success.

### Can AI visibility change without us doing anything?

Yes. On one site, ChatGPT referrals to pages nobody had touched grew 3.5 times from January to May 2026, which the authors put down to the growth of ChatGPT itself.

### How long should we measure before judging GEO results?

At least several weeks. One study found per-brand estimates took about 24 days of daily runs to become reasonably stable, and short windows mostly show noise.

### Can competitors lower our AI visibility?

Yes. Citation share is zero-sum within an answer, so competitors’ gains can lower your share even if your content and the engine did not change.

## Sources

- Martinez (2026), [Measuring GEO Visibility: Prompt Corpora Define the Answer Market](https://arxiv.org/abs/2609.06811), arXiv:2609.06811.
- Martinez (2026), [Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization (2023-2026)](https://arxiv.org/abs/2607.14035), arXiv:2607.14035.
- Watanabe and Nakayashiki (2026), [Disentangling Answer Engine Optimization from Platform Growth: A Log-Based Natural Experiment on ChatGPT Referral Traffic](https://arxiv.org/abs/2606.04362), arXiv:2606.04362.
- Sielinski (2026), [Quantifying Uncertainty in AI Visibility: A Statistical Framework for Generative Search Measurement](https://arxiv.org/abs/2603.08924), arXiv:2603.08924.
- Kumar (2026), [Generative Engine Optimization at Scale: Measuring Brand Visibility Across AI Search Engines](https://arxiv.org/abs/2606.20065), arXiv:2606.20065.
- Tannenbaum (2026), [Scoring With the Engine: Retrieval Exposure, Cross-Engine Divergence, and the Limits of Engine-Agnostic GEO Scores](https://arxiv.org/abs/2609.22655), arXiv:2609.22655.
- Schulte, Bleeker and Kaufmann (2026), [Don’t Measure Once: Measuring Visibility in AI Search (GEO)](https://arxiv.org/abs/2604.07585), arXiv:2604.07585.
- Underneath (2026), [Ask an AI the same question 5 times: do the brands change?](https://underneath.agency/research/ai-recommendation-consistency-study)

---

This is the Markdown twin of https://underneath.agency/resources/did-geo-work-raise-ai-visibility. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
