---
title: "Results: How We Measure GEO and AI Visibility | Underneath"
description: "How Underneath measures AI visibility in layers: what AI answers say, the evidence underneath, what people do and what the business gains. No invented scores."
canonical: "https://underneath.agency/results"
published: 2026-09-26
updated: 2026-09-27
publisher: "Underneath (https://underneath.agency/agent)"
entity: "https://underneath.agency/.well-known/entity.json"
---
How we measure

# A result is only proof *if you can check it.*

AI visibility is easy to make sound impressive: a percentage, a ranking, a screenshot of an answer. None of those means much without a baseline, a defined question, a consistent method and a connection to the business.

We measure the work in layers: **Visibility → Evidence → Behavior → Business.** That keeps three things apart: what changed in an AI answer, what changed in the information underneath it, and what changed for the company.

Each report shows what moved, what did not, what we believe caused it and what we would investigate next.

Visibility

## What does the market *see*?

We build a representative set of questions around the decisions your buyers make, not just searches for your brand name.

- Category questions
- Comparison questions
- Problem questions
- Product questions
- Location questions
- Alternative questions
- High-intent questions

We track how your business and competitors appear across those questions and how the answers change over time.

### We look at

Presence
:   Whether the business appears at all.

Position
:   Where and how prominently it appears when the answer involves multiple companies.

Description
:   What the system actually says about the business.

Recommendation
:   Whether the business is presented as an option when the question calls for one.

Citation
:   Which sources are being used to support the answer.

Persistence
:   Whether the pattern survives changes in wording and related questions.

**A single answer is an observation. A repeated pattern is a signal.**

Evidence

## What is underneath the *answer*?

Visibility tells us what happened. Evidence helps us understand why it may be happening.

For important answers, we trace the information underneath them. We look at the claims being made, the sources associated with those claims, the entities involved and the consistency of that information across the wider web.

### We investigate

Claims
:   What is being said about the company, product or service?

Sources
:   Where does the information come from?

Corroboration
:   Is important information supported beyond the company’s own website?

Freshness
:   Is the information current?

Competition
:   Where do competitors have a stronger or more complete information footprint?

Without this layer, a visibility change is hard to act on: you can see that something moved, but not what to do about it. Accuracy, consistency and coverage are measured in their own right, below.

Information health

## Accuracy is a measurement, *too*.

A business can be highly visible and still be badly represented.

AI answers can contain

- Outdated products
- Old pricing
- Incorrect locations
- Incomplete descriptions
- Outdated company information
- Incorrect relationships
- Unsupported claims
- Conflicting information

So we measure whether the facts that matter most are right.

### We track

Accuracy
:   Are important facts represented correctly?

Conflicts
:   Where do authoritative sources disagree?

Outdated claims
:   Which old information continues to surface?

Coverage
:   Which important facts are missing?

Consistency
:   Does the business describe itself consistently across important sources?

Appearing more often is only half the job. The other half is being described correctly when you do.

Behavior

## Did visibility change what people actually *did*?

An AI mention is not a business outcome. Wherever the data allows, we connect AI visibility to what people actually did.

### We measure

AI referrals
:   Sessions arriving from AI systems and AI-driven search experiences where attribution is available.

Engagement
:   What those visitors do after arriving.

Leads
:   Calls, forms, bookings, applications and other agreed conversion events.

Qualified demand
:   Which AI-referred inquiries become meaningful opportunities.

Revenue
:   Where the available attribution supports connecting AI-driven demand to revenue.

Assisted influence
:   Where AI or search visibility contributes to a journey without being the final recorded source.

We distinguish what we can directly attribute from what we can only reasonably observe.

Business

## The board doesn’t need an AI *score*.

It needs to know whether the work is paying off. Every engagement starts by agreeing what the business is trying to improve.

Depending on the company, that might be

- Qualified leads
- Pipeline
- Revenue
- Applications
- Bookings
- Cost per acquisition
- New customers
- Geographic demand
- Product adoption

AI visibility is then treated as one layer of the measurement system, not the final objective.

The question we answer is not whether a number on a dashboard went up. It is:

> Did the work produce a meaningful change in the outcomes we agreed to improve?

Measurement without false precision

## We don’t pretend the system is more observable than it *is*.

Models are updated, answers shift from one day to the next, sources come and go, and attribution is never complete. So we are careful about what the data can actually prove.

### We distinguish between

Observed
:   Something directly measured.

Correlated
:   Two things changed together, but causation is not established.

Attributed
:   The available tracking supports a connection between the change and the business outcome.

Hypothesized
:   A plausible explanation that needs further testing.

A report should tell you what we know and, just as plainly, what we don’t know yet.

Experiments

## Measure the intervention, not just the *outcome*.

When we make a meaningful change, we want to know why we made it and what we expect to learn. A typical cycle looks like this.

- ### Observation

  Something in the information environment stands out.
- ### Hypothesis

  We propose a reason it may be happening.
- ### Intervention

  We change the relevant part of the system.
- ### Retest

  We return to the relevant questions and measurements.
- ### Interpretation

  We determine what changed and how confident we should be in the explanation.
- ### Next decision

  We either build on the finding, revise the hypothesis or move somewhere else.

This prevents the program from becoming a collection of disconnected deliverables.

Our measurement model

## Four layers. One *picture*.

- ### Visibility

  **What is being said?** Mentions, recommendations, descriptions, citations and competitive presence.
- ### Evidence

  **What supports it?** Claims, sources, entities, corroboration, conflicts and information coverage.
- ### Behavior

  **What did people do?** AI referrals, engagement, inquiries, applications, bookings and conversions.
- ### Business

  **Did it matter?** Pipeline, customers, revenue, acquisition cost and the business metric that matters to you.

The layers should connect, but they should not be confused.

A visibility increase is not automatically a revenue increase. A citation is not automatically a lead. A lead is not automatically a customer.

We measure each step on its own terms.

Illustrative scenarios

## Different businesses require different *evidence*.

We don’t use the same dashboard for every client. What we measure follows the problem.

Scenario 01 · B2B software

### Assistants recommend three competitors when buyers ask for the best tools in the category.

Illustrative scenario, not a client result

Trace which comparison pages, reviews and analyst summaries the answers lean on, then earn a credible place in those sources instead of publishing more blog posts.

Share of AI answers mentioning the brand
:   Mentions

Citations of owned pages
:   Citations

Pipeline from AI referrals
:   Pipeline

- Generative engine optimization (GEO)
- Enterprise GEO
- What we would measure

Scenario 02 · Financial services

### Product pages sit below aggregators, and every new page needs compliance review.

Illustrative scenario, not a client result

Build an answer-first content model with a review workflow compliance can live with, and roll structured data out at template level.

Answers extracted by search engines
:   Answers

Product applications from AI referrals
:   Applications

Time from draft to approved page
:   Speed

- Enterprise GEO
- What we would measure

Scenario 03 · Enterprise software after a rebrand

### AI answers still describe discontinued products and old pricing, citing years-old community threads.

Illustrative scenario, not a client result

Correct the record at the source: clear, current brand facts on the site, and a disclosed expert presence in the communities models read.

Factual accuracy of AI answers
:   Accuracy

Outdated claims in answers
:   Claims

Community sentiment
:   Sentiment

- Community citations
- Generative engine optimization (GEO)
- What we would measure

Scenario 04 · Single-location service business

### Calls are flat, and AI assistants name three competitors when someone asks for a plumber nearby.

Illustrative scenario, not a client result

Fix call and form tracking first, so every lead has a source. Then correct the Google Business Profile and location-page facts, build them around the reasons customers give in reviews, and work toward being the business local AI answers name for the services that turn into booked jobs.

Tracked calls and form fills
:   Calls

Cost per booked job
:   Cost

Local AI answers naming the business for core services
:   AI answers

- Local GEO
- Google Business Profile
- What we would measure

What goes in front of your board

## A report should make decisions *easier*.

The cadence depends on the engagement. The questions don’t.

Weekly

- What shipped?
- What changed?
- What did we learn?
- What happens next?

Monthly

- How are the agreed metrics moving?
- Which experiments are producing useful evidence?
- What should change in the work?

Quarterly

- What changed in the information environment?
- What changed in visibility?
- What changed in customer behavior?
- What changed in the business?
- What did we learn that changes the next quarter?

The objective is not a prettier dashboard. It is a clearer decision.

What we won’t do

## We don’t publish numbers we can’t *defend*.

- No invented “AI visibility score.”
- No impressive percentage without a baseline.
- No case-study language around illustrative scenarios.
- No claim that an AI mention caused revenue simply because the two appeared in the same quarter.
- No screenshot presented as proof of a durable result.

And no client result is published until the client has approved it. Until then, every scenario on this site is labeled illustrative.

The point of measurement

## Make the invisible *discussable*.

AI visibility feels hard to measure because every answer is generated on the spot and the systems behind it keep changing. That doesn’t make measurement impossible. It makes discipline essential.

1. Define the question.
2. Establish the baseline.
3. Observe the answer.
4. Trace the evidence.
5. Measure the behavior.
6. Connect it to the business.
7. State what is known.
8. State what isn’t.

Then decide what to do next.

Free strategy call

## Your baseline is the first result worth knowing.

On a free 30-minute call we take a first look at how you show up in search and AI answers and what your site gives them to cite. Then we tell you plainly whether a full audit is worth it.

[See where you stand](https://underneath.agency/contact)[How we work](https://underneath.agency/approach)

---

This is the Markdown twin of https://underneath.agency/results. The HTML page is canonical. Publisher: Underneath, https://underneath.agency/agent. Site index: https://underneath.agency/llms.txt.
